Strix Halo NPU finally put to work: local 125B MoE replaces 95% of cloud coding agent calls
stereohype · reddit · 2026-10-08
A Reddit user wired Halogen's NPU endpoints into the pi coding agent on a 70W Strix Halo tablet, running local Qwen3.8 Flash-Next (125B MoE, 64 tok/s decode, 0.03s first token) and claims it now replaces 95% of cloud model calls, with only the hardest tasks still going to GLM 5.3 or Opus.
Key numbers:
- Same bug fix: 13.6 min with NPU search vs 18.7 min without
- Seven-model bake-off on one real timeshift error: Flash-Next delivered the full verified answer fastest (2m55s, tracing a notify-send race to an upstream PR), beating max-tier 753B GLM 5.3 (6m15s)
- NPU handles four sidecar jobs: file search (0.1s/lookup, 15/20 vs ripgrep's 9/20), duplicate scan (45 pairs in 8.4s), yes/no decision routing (0.8B, 120ms, 78% accurate), prompt-injection screening (0.7s/msg, 42% recall, zero false positives)
- Compaction of a 194k session: 50s via sidecar vs 166s on the main model, 97% cache hit
Honest caveat: the GPU still does the thinking — the NPU changed which tokens get spent, not raw speed. Fully local with 262k context; repo and full comparison docs are linked.
More from coding & agent
- Agent Guard adds guardrails to running Claude Code in YOLO mode — Arindam_1729 · 2026-10-09
- OpenClaw slashes 1.3M flaky tests, doubling CI speed in big cleanup push — heyneighbor · 2026-10-09
- How do you decide what to optimize first in production AI agents? — Successful-Ask736 · 2026-10-09
- Matt Pocock's "Skills for Real Engineers" repo hits 280k GitHub stars — adnan_hashmi · 2026-10-09
- Steve Yegge: Rex Alpha Makes 20+ Claude Code Sessions Feel Local From a Hotel Laptop — Steve_Yegge · 2026-10-09
- Agent harnesses are crutches; least-privilege action authorization is what matters — andreisavu · 2026-10-09