Community finds ~17% of MiMo-V2.6-Pro's ~1T params can be pruned with no KLD change

QuixiAI · x · 2026-09-29

User turboderp reports that MiMo-V2.6-Pro (1T params) has highly redundant weights: removing expert blocks from layers 1-12 and attention blocks from layers 1-4, 6 and 8-11 leaves end-to-end KDL largely unchanged, with contributions at BF16 rounding-error scale. That suggests 17% of the parameters — six whole Qwens' worth — can be trimmed with the model essentially intact. The finding is unconfirmed and the author is seeking independent verification; if it holds, it's a notable signal about MoE parameter efficiency.

Original post →

More from Infra

Infra channel →