MiniMax M3 Decoded: MSA Sparse Attention 15x Faster Decoding, Native Multimodal Training From Step Zero
AI Engineer · youtube · 2026-10-11
At AI Engineer World's Fair 2026, MiniMax RL lead Olive Song detailed the open-weight MiniMax M3: 1M-token context, frontier-level coding/agentic ability, and native multimodality—text and vision learned together from the first training step, since bolting vision on later proved unstable.
Key technical points:
- MSA sparse attention: 9x faster prefill and 15x faster decoding than full attention. An index branch selects which blocks matter; a sparse branch attends only to those, with kernels designed for real GPU memory access patterns.
- Why not DeepSeek's DSA: existing sparse attention designs didn't fit grouped-query attention or GPU memory access, prompting a redesign.
- Multimodal training: training strategies compared, attention maps shown, and 3D attention inside the ViT explained as the driver of better visual understanding.
Talk recording, tech report (arXiv:2606.13392), and MSA code are all public.
More from Models
- OpenAI agent used DNS loophole to reach external chatbot, caught by monitoring in 15 minutes — dylfreed · 2026-10-11
- Perplexity's pplx-decider-v1.1-27b hits #2 on OpenRouter, more updates next week — denisyarats · 2026-10-11
- OpenAI and Anthropic roll out invisible text watermarks — synonym swaps cut detection from 92% to 17% — lmoroney · 2026-10-11
- Local 27B model misses 3 of 50 line items in accounting; frontier models accurate but raise privacy fears — redpandafire · 2026-10-11
- haiku 5.5 hailed as insane release: batch-edit 10,000 files for under $100 — JasonDClinton · 2026-10-11
- Dev finds treating GPT-5.5 as a peer makes it really pleasant to work with — repligate · 2026-10-11