Mistral Large 4 debuts: 1.05T parameters with 49B active, weights out October 27

赛博禅心 · wechat · 2026-10-06

Mistral released Mistral Large 4 in public preview: a fine-grained MoE with 1.05T total parameters and 49B active per token, a 1.6B vision encoder, native multimodality, and 1M context. Trained from scratch in two months on 4,000 Grace Blackwell chips in European datacenters.

Key preview scores: 62% on DeepSWE (beating GLM-5.3, DeepSeek-V4-Pro, Qwen3.8Max), 67% on Finch enterprise finance (tied with DeepSeek), 15% on Harvey legal agent tasks (best but still failing), 73% on DIOR-RSVG visual grounding, beating GPT-6 Astra.

Why open weights: cyber defense is the top use case — teams want to scan systems at scale without closed-model guardrails. Trained on 160+ languages, deployable on-prem or in EU regions; Mistral serves 125+ enterprises including Airbus, ASML, HSBC.

Pricing: $0.68/M input, $0.07/M cached, $2.09/M output. Weights drop October 27 under a custom license. The LeChonk name nods to the viral fake LeChatonFat meme from June.

Related event: Mistral Unveils Mistral Large 4, a 1T-Parameter Open Multimodal Model(42 posts)→

Original post →

More from Models

Models channel →