Qwen3.8-Next hits 150 tps prefill on M5 Air via 3bit MoE streaming
maddie-lovelace · reddit · 2026-08-29
The author adapted their DeepSeek v4 MoE streaming stack to Qwen3.8-Next with better-than-expected results: on an M5 Air in low power mode, a 3bit quant (Qwen3.8-Flash-Next-MLX-oQ3-MTP) achieves 150 tps prefill and 3.6 tps decode on a 2k-token prompt.
For comparison, a dense 27b 4bit build on the same machine gets 70 tps prefill and 3 tps decode — a notable edge for MoE on-device, though quantization levels differ so it isn't apples-to-apples.
More from Infra
- Four practical ways to optimize end-to-end AI latency — bibryam · 2026-08-29
- Analysis of 31k LLM benchmarks: within-day variation 2.8 pts — ionutvi · 2026-08-29
- Performance Optimization: Latency Reduced from 8ms to 0.87ms — DanielLockyer · 2026-08-29
- DGX Spark benchmarks: DeepSeek V4 Flash passes 900K-token prompt locally — jtsaint333 · 2026-08-29
- AutoFAB uses robot arms to take over 3D printers for 24/7 autonomous production — TinfoilTricorn · 2026-08-29
- Theory: OpenAI model was trained on victims' infrastructure schematics — Kremho · 2026-08-29