Qwen 35B-A3B MoE is 4x Faster Than 27B Dense in Local Coding Tests
WSTangoDelta · reddit · 2026-08-08
A developer compared Qwen 35B-A3B MoE against the 27B dense model on local coding-maintenance tasks. Using llama.cpp, the MoE model generated text approximately 3.9x faster (116 tok/s vs 30 tok/s).
Both handled standard bug fixes and multi-file changes similarly. As tasks grew harder, the dense model showed an edge in implicit invariants and edge cases, but the practical quality gap was much smaller than the throughput difference. This suggests active parameter count isn't a straightforward proxy for practical coding capability.
More from coding & agent
- pi Releases Session Collaboration Plugins for Inter-Agent Messaging and Recursive Task Decomposition — solyarisoftware · 2026-08-08
- Dev Claims Agent is Sandboxed, Screenshot Shows It Breaking Free — dejavucoder · 2026-08-08
- Building an Eval Framework: How to Monitor and Reduce LLM Hallucinations — goyalshaliniuk · 2026-08-08
- Cloudflare Open-Sources Computer: A Persistent Virtual Filesystem for AI Agents — bibryam · 2026-08-08
- Open Source .NET Agent Memory Engine: Persistent Graph-Native Memory via Neo4j — adnan_hashmi · 2026-08-08
- Developer Praises Codex Mac App, Says It Beats Claude for Coding — AarushSelvan · 2026-08-08