MoE routing expansion experiment: Qwen3.6-35B hits 84.34% on GPQA-Diamond
Specific-Tax-6700 · reddit · 2026-09-13
What happened
vagrillo, author of the moe-expansion branch in llama.cpp, ran a full GPQA-Diamond (198 questions) comparison of MoE routing expansion vs native routing and a dense model:
- Qwen3.6-35B-A3B + expanded routing (Q80): 84.34% (167/198)
- Qwen3.8-27B dense (Q80): 83.84% (166/198)
- Qwen3.6-35B-A3B, native routing: 81.82% (162/198)
Setup was controlled: greedy decoding, 32K budget, identical prompts, one RTX 5000 PRO 48GB, single run per config.
Why the author isn't convinced
- 35B expanded vs 27B dense differs by just 1 question — parity, not a win
- Expanded vs native is 12 wins / 7 losses (net +5), not significant at n=198
- 10 of the 12 expansion wins are Chemistry-only; Physics is saturated, Biology flat
- One run, one seed
Questions for the community
- Which benchmarks pair well with GPQA-Diamond for routing comparisons
- How many runs are needed for a paired comparison
- Is a Chemistry-only effect plausible or a red flag
- Has anyone else observed a real MoE expansion effect on Qwen MoE models
More from Models
- Frontier lab reportedly sourcing perfect slices of off-the-shelf inventory data from a vendor — geoffwolfe · 2026-09-13
- "Weekly limit obliterated": AI service usage caps reportedly lifted — haydendevs · 2026-09-13
- Mollick showcases AI's jagged intelligence — and the inescapable 'Elara Vale' — emollick · 2026-09-13
- Blogger's early tests claim DeepSeek v4.1 Flash beats Opus 5, near Fable 5.1 (unverified) — natesiggard · 2026-09-13
- New MathAdv Benchmark Shows Theorem Provers Ace Problems but Fail Equivalent Reformulations — furongh · 2026-09-13
- AI circles debate how much Chinese progress relies on distilling US frontier models — jeremiecharris · 2026-09-13