121M on-device model hits 86.0 on tool calling, nears DeepSeek V4 Flash's 88.4
Henrie_the_dreamer · reddit · 2026-09-18
Cactus Compute open-sourced Needle 3 (29–121M params), a foundation model purpose-built for on-device automation: tool calls, structured extraction and embeddings — no chatting, unsupported requests return an empty list. Architecture: dense FFNs replaced by a Monarch Hadamard MLP (25.6K params/layer), with knowledge stored in "engram" hashed n-gram tables that cost no arithmetic — 70.8M of 121M params live there, so the model computes like a 50M one (100 MFLOPs/token vs 296). Arguments are spans of the request, emitted under a byte-level grammar compiled from your schema, so JSON always parses. On Mobile Actions (961 phone commands, exact match), the 121M model scores 86.0 at 2-bit, beating LFM2.5 1.2B (82.4), Qwen3.5 0.8B (76.0) and Apple's 3B on-device FM (57.6), nearing cloud DeepSeek V4 Flash (88.4), while trailing on general suites like BFCL v4 and DSTC8. Available on Hugging Face, GitHub, PyPI, with a browser sandbox.
Related event: Cactus Releases Needle 3 On-Device Model, Hits 86 on Function Calling(2 posts)→
More from Models
- ChatGPT co-inventor launches Jev: claims 20-200x faster, 40-400x cheaper than LLMs — GabGarrett · 2026-09-18
- Ran Jev across 10 services for hours, still couldn't spend $1 — multiply_matrix · 2026-09-18
- ChatGPT co-inventor launches Jev, claiming 200x faster, 400x cheaper frontier model — multiply_matrix · 2026-09-18
- Tencent's Hy4 Preview ranks #4 among open-weight models, cheapest in top ten — mariofilhoml · 2026-09-18
- Simple letter-counting test exposes huge gap: GPT-6-Astra hits 93%, Fable 5.1 flounders — scaling01 · 2026-09-18
- Jev, a 'System One' model by Typesafe, launches on OpenRouter with typed decisions instead of text — majidmanzarpour · 2026-09-18