Xiaomi Previews MiMo-V3's HySparse2 Architecture, Cutting Million-Token Prefill Compute ~5x

量子位 · wechat · 2026-09-24

Right after MiMo-V2.6 shipped, Xiaomi's Luo Fuli previewed MiMo-V3's new core architecture, HySparse2, built for the agent workload where tool calls emit dozens of tokens but return tens of thousands:

Results at 80B total / 3B active params: at 1M context, prefill compute is 1/5.02 of HybridSWA, and FP8 KVCache drops from 12.09GB to 2.69GB (4.5x). At 256k context, RULER-v2 hits 58.45 vs 32.61 (HySparse) and 35.74 (HybridSWA); average MRCR-v2/RULER-v2 gains over HySparse are +11.30 and +19.81 points, with general knowledge/code roughly on par. Evaluations extend to 256k; real-world latency awaits MiMo-V3's release.

Related event: Xiaomi Reveals HySparse2 Architecture for MiMo-V3(5 posts)→

Original post →

More from Models

Models channel →