CPU-Only LLM Tests: 35B MoE at Q2 Beats a 2B Model Despite Half the Speed
ML-Future · reddit · 2026-09-08
The author debates the future of local LLMs — tiny-but-smart vs. huge-but-optimized — with hands-on CPU-only benchmarks.
- MiniCPM5 2B Q8: 6 t/s without vision, but makes many mistakes
- Qwen3.6 35B (MoE) Q2XXS: 3 t/s but far more impressive, usable results
- Verdict leans toward heavily-quantized MoE large models over small dense ones for CPU-only setups
Follows up the author's earlier post on running Qwen3.6 35B without a GPU.
More from Infra
- Yacine urges building sovereign AI infrastructure or losing privacy and IP — yacineMTB · 2026-09-08
- Crusoe raises $3B at $30B valuation as it builds Stargate, delivered 200MW in 11 months — 快鲤鱼 · 2026-09-08
- HydraDB: a Rust graph database that lives entirely in S3, with nothing on disk — thisdudelikesAI · 2026-09-08
- Running Qwen 27B and Gemma 31B locally on one RTX 4090: quantization and context tradeoffs — MooseEfficient2151 · 2026-09-08
- Could distributed iPhones form an inference network? A 2AM open question — gajesh · 2026-09-08
- Candidate depth 10→500 lifts score by just 0.01: Qdrant on diagnosing before tuning vector search — qdrant_engine · 2026-09-08