Qwen3.8 Expert Analysis: 50% Can Be Pruned With Minimal Loss
EyalToledano · x · 2026-08-27
I scored all 24,576 experts in Qwen3.8-Flash-Next using 700k tokens of session traces. Running in 4-bit on an M4 Max, I found that at the median layer, 64 of 512 experts carry half the routing mass, and 256 carry 91%. This implies half the experts are effectively passengers, suggesting we could likely prune 50% of experts while retaining 90%+ accuracy.
Related event: Analysis Finds Half of Qwen3.8 Experts Nearly Idle(2 posts)→
More from Research
- TIDES Dataset: Longitudinal Bilingual Record of 12 Teams' Collaboration — josephseering · 2026-08-27
- Kyoto U's MemUse: Natural Integration Outperforms QA in Evaluating Conversational Memory — Kyoto-University · 2026-08-27
- GPT-5.6 Builds New Kernel, Achieving 9.7x Speedup on TPU — HuaxiuYaoML · 2026-08-27
- RSI-Exam Benchmark Launches to Test AI Recursive Self-Improvement — HuaxiuYaoML · 2026-08-27
- Gordian Screens 1,327 Targets In Vivo, Accelerating Drug Discovery — juanbenet · 2026-08-27
- Discussion on Multi-Agent Reward Schemes and Convergence — jessi_cata · 2026-08-27