Show HN: Cactus Needle 3 — 8-29MB on-device models claim DeepSeek V4 Flash-level automation

HenryNdubuaku · hn · 2026-09-18

Cactus releases Needle 3 on HN: tiny automation-focused models (tool calls & structured JSON only, no chat) with 2-20 deployable subnetwork layers (25M-121M params at 2-bit) shipping as 8-29MB binaries. A Monarch Hadamard MLP replaces the dense FFN at O(d√d) cost. On a Raspberry Pi 5 it decodes up to 4k tokens/s. On Mobile Actions the 20-layer model scores 86.0 vs LFM2.5 1.2B (82.4), Qwen3.5 0.8B (76.0) and Apple's on-device model (57.6). Ships in 7 languages; fine-tuned 4-layer versions reportedly match DeepSeek V4 Flash on narrow tasks.

Original post →

More from Infra

Infra channel →