Mistral releases Large 4 'Le Chonk': 1.05T-param MoE, 49B active, 1M context, 3,800 Blackwell GPUs
_AndrewZhao · x · 2026-10-07
- Mistral AI released Mistral Large 4 (nicknamed Le Chonk) in public preview: a granular MoE with 1.05T total parameters, 49B active per token (4.7% of weights), a 1.6B vision encoder, and a 1M-token context window with native image input.
- Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters; weights ship at end of October, so self-hosting isn't possible yet.
- API is live as mistral-large-4 at $1.36/1M input and $4.18/1M output tokens.
- Benchmarks: 93% Cybench, 82% CyberGym-E2E (top 5 on Artificial Analysis Cyber Index); 61.7% DeepSWE v1.1, 28.3% Terminal-Bench 4.0, 49.8% Coding Agent Index.
- Deployment note: all 1.05T weights must sit in memory — plan hardware around total, not active, parameters.
Related event: Mistral Launches Mistral Large 4, a 1T-Parameter Open-Weight Flagship(67 posts)→
More from Models
- Early verdict on Meta Muse: "not very good" — HanchungLee · 2026-10-07
- Claude Surprises User by Offering to Switch to Work Mode Mid-Maxscript Coding — Efistoffeles · 2026-10-07
- DIY eval: pplx-decider-1.1 hits 95.6% agreement with a frontier model — bo_wangbo · 2026-10-07
- Google's EmbeddingGemma 2 (740M) claims to beat embedding rivals twice its size — The Decoder · 2026-10-07
- Mistral Large 4 launches as CEO claims it beats Chinese models on cyber capabilities — BLUECOW009 · 2026-10-07
- Florida woman charged with felony after Claude flagged her threat and a human reviewer tipped police — XFreeze · 2026-10-07