Cactus Releases Needle 3: An 8-29MB Foundation Model Running 4k tokens/s on a Raspberry Pi 5
airesearch12 · x · 2026-09-18
Cactus released Needle 3, a foundation model for mobile, wearables, robots, smart home, automotive and microcontrollers. The whole model is a single 8-29MB binary built on its Simple Attention Network; one set of weights slices from 2 to 20 layers (25-121M parameters at CQ2-bit), and it runs locally at up to 4k tokens/s decode on a Raspberry Pi 5.
Key design: no chatting, pure automation.
- Tool calling: picks the right functions and fills arguments from user speech; uncovered requests return an empty list, not a guess
- Structured extraction: schema in, typed fields out, with decode grammar guaranteeing parseable output
- Text embeddings from the same model for local search, matching and routing
Cactus claims the 121M version, trained on 360B tokens of structured data, beats models 10x its size on mobile tool calls and matches 2-3x bigger models on extraction; a fine-tuned variant passes DeepSeek V4 Flash from 4 layers up.
Related event: Cactus Releases Needle 3 On-Device Model, Hits 86 on Function Calling(2 posts)→
More from Infra
- 600 tok/s single-request on Qwen 35B with Ninfer on an RTX Pro 6000 — CharlesStross · 2026-09-18
- Cadence sees India's EDA market doubling to $7.82B by 2031 — bookwormengr · 2026-09-18
- Google Open-Sources Agent Substrate on GKE: 10x Density, 1,000+ Dormant Agents per Host — blaizedsouza · 2026-09-18
- Redditor crams six V100 GPUs into a standard full-tower case for local LLM inference — Odd_Caterpillar_2994 · 2026-09-18
- Crusoe raises $3.9B at $30.9B valuation to build data centers and modular AI factories — TechCrunch AI · 2026-09-18
- A 2.5-hour first-principles primer on the semiconductor supply chain worth your time — blaizedsouza · 2026-09-18