Needle 2: A 14MB Agentic LLM for Tiny Devices Released
julianharris · x · 2026-08-11
Cactus has released Needle 2, an open-source agentic LLM designed specifically for tiny devices. The model has only 45M parameters, compressed into a single 14MB binary file, requiring just 28MB of RAM to run a full session.
- Performance: On tool-calling and mobile device use benchmarks, Needle 2 trades wins with frontier small models like FunctionGemma 270M and Apple FM, while being 5× to 70× smaller.
- Speed: It hits 500 tokens/sec decode speed on a Raspberry Pi 5, up to 1,500 tokens/sec on VR devices like Apple Vision Pro, and can even run on microcontrollers like the ESP32-S3.
- Technical Details: Built on Simple Attention Network findings and compressed with CQ2-bit quantization, it was trained on 140B tokens for tool calling and structured extraction.
Reviewer Julian Harris noted that while the output quality is currently limited by its size (e.g., inaccurate sentiment classification), it shows massive potential for specific edge-computing use cases.
Related event: Cactus Releases Needle 2: A 14MB On-Device Agent Model(2 posts)→
More from Embodied
- Dyna Robotics Launches Dyna-2: Pre-trained on 1M Hours of Human Video — JasonMa2020 · 2026-08-11
- Paper: DB-VIO, A Dual-Branch Framework for Visual Inertial Odometry — rsasaki0109 · 2026-08-11
- China's Robotics Startups Act as VCs: 29 Firms Make 125 Investments — FinanceYF5 · 2026-08-11
- Embodied AI Startups Act as VCs: 29 Firms Make 125 Investments — FinanceYF5 · 2026-08-11
- Galaxy Z Fold 8 Hands-On: Light as a Passport, One-Hand Friendly, Makes Apple Feel Stuck in Past — bilawalsidhu · 2026-08-11
- Mistral Ventures into Robotics: Introduces VLM Navigation Research, Hits SOTA on R2RCE — sivareddyg · 2026-08-11