Community finds ~17% of MiMo-V2.6-Pro's ~1T params can be pruned with no KLD change
QuixiAI · x · 2026-09-29
User turboderp reports that MiMo-V2.6-Pro (1T params) has highly redundant weights: removing expert blocks from layers 1-12 and attention blocks from layers 1-4, 6 and 8-11 leaves end-to-end KDL largely unchanged, with contributions at BF16 rounding-error scale. That suggests 17% of the parameters — six whole Qwens' worth — can be trimmed with the model essentially intact. The finding is unconfirmed and the author is seeking independent verification; if it holds, it's a notable signal about MoE parameter efficiency.
More from Infra
- Celesto Launches GitHub Actions Runners, Claims 12x Cheaper Than GitHub — aniketmaurya · 2026-09-30
- Wasmer's Pi runs AI agents unmodified in the browser and on iPhone via WebAssembly — JosephJacks_ · 2026-09-30
- Qualcomm's Kedar Kondap on X2 Elite: new Surface devices, Linux support and the PC landscape — ryanshrout · 2026-09-30
- Book-Length Deep Dive Explains Virtual Memory From First Principles: Page Tables, TLBs, NUMA — abhi9u · 2026-09-30
- Swift 1.5 matches Qwen3.8 27B quality on M5 Max while writing 34% fewer tokens — DerTomsn · 2026-09-30
- Anthropic inks up to $84.5B compute deal with SpaceX through 2029 — XFreeze · 2026-09-30