MiniMax M3 Decoded: MSA Sparse Attention 15x Faster Decoding, Native Multimodal Training From Step Zero

AI Engineer · youtube · 2026-10-11

At AI Engineer World's Fair 2026, MiniMax RL lead Olive Song detailed the open-weight MiniMax M3: 1M-token context, frontier-level coding/agentic ability, and native multimodality—text and vision learned together from the first training step, since bolting vision on later proved unstable.

Key technical points:

Talk recording, tech report (arXiv:2606.13392), and MSA code are all public.

Original post →

More from Models

Models channel →