CPU-Only LLM Tests: 35B MoE at Q2 Beats a 2B Model Despite Half the Speed

ML-Future · reddit · 2026-09-08

The author debates the future of local LLMs — tiny-but-smart vs. huge-but-optimized — with hands-on CPU-only benchmarks.

Follows up the author's earlier post on running Qwen3.6 35B without a GPU.

Original post →

More from Infra

Infra channel →