MiniMax M3: intern-designed sparse attention powers 1M-token agent context
AI Engineer · youtube · 2026-09-04
Hugging Face co-founder Thomas Wolf and MiniMax RL lead Olive Song unpack MiniMax M3 at AI Engineer World's Fair: a functional 1M-token context window combined with coding, agentic and multimodal capabilities, built because short context windows can't support agents working across long conversations and tool outputs. Highlights: the sparse-attention architecture was designed by an intern, M3 was trained multimodal from the first step, MiniMax serves 300M users while continuing to open-source frontier models, and M3 is already helping the team build M3.1—with multi-agent systems named as the next frontier.
More from Models
- After a major model upgrade, prune your AGENTS.md: Astra follows stale rules too literally — keyanzhang · 2026-09-05
- Grok web gets a cleaner redesign with a refreshed UI — XFreeze · 2026-09-05
- eyebench author says no v4, moving on to harder benchmarks — adonis_singh · 2026-09-05
- Astra-max claims vastly better intelligence-per-token even at low reasoning — adonis_singh · 2026-09-05
- Astra-max hits 95% on eyebench-v3 at half the cost of Sol-max, tokens ~3.8x fewer — adonis_singh · 2026-09-05
- pass@1 dead even, but Fable 5.1 wins pass@k over GPT 6 Astra — zainhas · 2026-09-05