Show HN: Cactus Needle 3 — 8-29MB on-device models claim DeepSeek V4 Flash-level automation
HenryNdubuaku · hn · 2026-09-18
Cactus releases Needle 3 on HN: tiny automation-focused models (tool calls & structured JSON only, no chat) with 2-20 deployable subnetwork layers (25M-121M params at 2-bit) shipping as 8-29MB binaries. A Monarch Hadamard MLP replaces the dense FFN at O(d√d) cost. On a Raspberry Pi 5 it decodes up to 4k tokens/s. On Mobile Actions the 20-layer model scores 86.0 vs LFM2.5 1.2B (82.4), Qwen3.5 0.8B (76.0) and Apple's on-device model (57.6). Ships in 7 languages; fine-tuned 4-layer versions reportedly match DeepSeek V4 Flash on narrow tasks.
More from Infra
- Google engineers: LLM benchmark harnesses silently drop requests — 200 QPS in, 38 out — AI Engineer · 2026-09-20
- The rig built to run Emacs and doomscroll X is now worth more than its owner's car — tetsuoai · 2026-09-19
- Apple M4 sustains 10 instructions per cycle, beating most rivals; M5 speedup explained — lemire · 2026-09-19
- Apple M6 bumps cores to 12 with two super cores; CPUs keep improving fast — lemire · 2026-09-19
- Apple M-series chips gained ~50% Geekbench 6 performance over three years — lemire · 2026-09-19
- Inside OpenAI's inference routing: why the proportional controller had to go — AI Engineer · 2026-09-19