Needle 2 - The 14 MB Agentic LLM for Tiny Devices | Cactus
- Today we release Needle 2: an open 45M-parameter model for tool calling, device use and structured extraction.
- The whole model is a single 14MB binary that runs a full session in 28MB of RAM.
- It is built on our Simple Attention Network findings, compressed to CQ2-bit with Cactus Quants, and baked into its own engine.
Unverified
- Today we release Needle 2: an open 45M-parameter model for tool calling, device use and structured extraction.
- The whole model is a single 14MB binary that runs a full session in 28MB of RAM.
- It is built on our Simple Attention Network findings, compressed to CQ2-bit with Cactus Quants, and baked into its own engine.
Sources: Cactuscompute