Training-free MoE router tweak cuts Qwen 35B reasoning tokens by 8.5%
Specific-Tax-6700 · reddit · 2026-09-04
A new paper describes a training-free inference-time MoE optimization: expanding the expert budget (A3B → A4B+) only in late transformer layers with linear decay. On full MMLU-Pro (714 questions, Qwen3.6-35B-A3B) it cuts mean reasoning tokens 8.5% and latency 10.9% while accuracy stays statistically unchanged (84.5% vs 84.0%). Dubbed "Succinct Convergence," the method needs zero training; paper, beta code and full JSON results are on Zenodo/GitHub.
More from Research
- De novo designed antibodies protect against lethal cobra venom in vivo — BrianHie · 2026-09-04
- Roman Telescope: 100x Hubble's field of view, to catalog 1 billion galaxies in year one — PeterDiamandis · 2026-09-04
- New paper: two tales of the geometric Jensen-Shannon divergence as JSD regularizations — FrnkNlsn · 2026-09-04
- Eric Topol's Lancet essay reviews AI health models that predict disease years ahead — EricTopol · 2026-09-04
- OpenCLIP's new native ModernText encoder: more customizable than Transformers ModernBERT, decoder-capable — wightmanr · 2026-09-04
- Johns Hopkins study shows LLMs fail to ignore training-time knowledge when context conflicts — mdredze · 2026-09-04