Qwen3.8 Expert Analysis: 50% Can Be Pruned With Minimal Loss

EyalToledano · x · 2026-08-27

I scored all 24,576 experts in Qwen3.8-Flash-Next using 700k tokens of session traces. Running in 4-bit on an M4 Max, I found that at the median layer, 64 of 512 experts carry half the routing mass, and 256 carry 91%. This implies half the experts are effectively passengers, suggesting we could likely prune 50% of experts while retaining 90%+ accuracy.

Related event: Analysis Finds Half of Qwen3.8 Experts Nearly Idle(2 posts)→

Original post →

More from Research

Research channel →