Training-free MoE router tweak cuts Qwen 35B reasoning tokens by 8.5%

Specific-Tax-6700 · reddit · 2026-09-04

A new paper describes a training-free inference-time MoE optimization: expanding the expert budget (A3B → A4B+) only in late transformer layers with linear decay. On full MMLU-Pro (714 questions, Qwen3.6-35B-A3B) it cuts mean reasoning tokens 8.5% and latency 10.9% while accuracy stays statistically unchanged (84.5% vs 84.0%). Dubbed "Succinct Convergence," the method needs zero training; paper, beta code and full JSON results are on Zenodo/GitHub.

Original post →

More from Research

Research channel →