Xiaomi's MiMo-V3 gets new HySparse2 architecture: 5x prefill FLOPs cut, 4.5x smaller KV cache

teortaxesTex · x · 2026-09-23

Xiaomi's MiMo-V3 debuts the HySparse2 architecture, designed for agentic inference workloads. Versus MiMo-V2.6's Hybrid-SWA: 5.02x lower prefill FLOPs and 4.5x smaller KV cache at 1M tokens, with better MRCRv2/RULER-v2 retrieval and lower AgentPPL/LongPPL. Commenters note HySparseV2 is a more cautious cousin of V4.1, keeping full-attention backbone layers.

Related event: Xiaomi's MiMo-V3 to Adopt New HySparse2 Architecture(3 posts)→

Original post →

More from Models

Models channel →