Mistral Large 4 preview: 1T params, 49B active, trained on 3,800 Grace Blackwell GPUs
Simon Willison · rss · 2026-10-07
Simon Willison reviews the Mistral Large 4 preview: a 1-trillion-parameter model with 49B active parameters, trained on Mistral's own cluster of 3,800 NVIDIA Grace Blackwell GPUs. Available via API now, with open weights promised by end of the month.
Key points:
- Only two reasoning levels (none/high); notably the 'high' pelican looked better while using just 2,717 output tokens vs 3,275 for 'none'
- Scores 38 on Artificial Analysis, just behind the 552B DeepSeek 4.1 Flash — a huge jump from Mistral Large 3's score of 9 last December
- Willison: not Fable-class, but Mistral is back to roughly 6 months behind the frontier
More from Models
- OpenAI shipped Decisions API just two weeks after first prototype — stevenheidel · 2026-10-07
- User marvels as ChatGPT shows humanlike academic skills — dioscuri · 2026-10-07
- Why Hybrid Models May Scale Better Downstream: The Inductive-Bias Argument — kalomaze · 2026-10-07
- Nous Research launches Hermes Index agent leaderboard, Claude Opus 5.5 tops at 63.31 — NousResearch · 2026-10-07
- Many Mistral Large 4 failures traced to reasoning mode not being enabled — qtnx_ · 2026-10-07
- Early Opus 5.5 user says hype is overblown: shortcuts, wrong assumptions, sloppy work — haider1 · 2026-10-07