MoE inference study says dropping 28% of routed experts left GSM8K unchanged
dai_app · reddit · 2026-07-25
A small MoE inference study found that the routing tail is surprisingly redundant.
- On a 35B MoE with top-8-of-256 routing, dropping about 28% of the lowest-weight routed experts did not change GSM8K accuracy.
- Baseline was 12/15 correct; every threshold tested, including the most aggressive one, was 13/15.
- 12 of 15 questions produced the exact same final answer across all settings, and response lengths stayed flat.
The setup used greedy decoding, so differences came only from skipping experts. The author notes the sample is tiny and the result is not a general proof, but it suggests the router’s low-weight tail may be a useful target for inference speedups when memory traffic is the bottleneck.
More from Infra
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- MiniBot 2.40 adds xAI, HF Studio and vLLM support with inline media tools — Creative-Type9411 · 2026-07-27
- Apple smart glasses, Nvidia-SK AI data center deal, and Ctrip’s RMB 5.179 billion fine headline a tech roundup — APPSO · 2026-07-27
- DeepSeek funding rumor, EU AI transparency rules and OpenAI agent incident make a packed AI news roundup — 创业邦 · 2026-07-27
- QuixiCore argues native quantized kernels beat dequant-then-generic execution — QuixiAI · 2026-07-27