MoE routing expansion experiment: Qwen3.6-35B hits 84.34% on GPQA-Diamond

Specific-Tax-6700 · reddit · 2026-09-13

What happened

vagrillo, author of the moe-expansion branch in llama.cpp, ran a full GPQA-Diamond (198 questions) comparison of MoE routing expansion vs native routing and a dense model:

Setup was controlled: greedy decoding, 32K budget, identical prompts, one RTX 5000 PRO 48GB, single run per config.

Why the author isn't convinced

Questions for the community

Original post →

More from Models

Models channel →