Strix Halo NPU finally put to work: local 125B MoE replaces 95% of cloud coding agent calls

stereohype · reddit · 2026-10-08

A Reddit user wired Halogen's NPU endpoints into the pi coding agent on a 70W Strix Halo tablet, running local Qwen3.8 Flash-Next (125B MoE, 64 tok/s decode, 0.03s first token) and claims it now replaces 95% of cloud model calls, with only the hardest tasks still going to GLM 5.3 or Opus.

Key numbers:

Honest caveat: the GPU still does the thinking — the NPU changed which tokens get spent, not raw speed. Fully local with 262k context; repo and full comparison docs are linked.

Original post →

More from coding & agent

coding & agent channel →