FULL STORY

Mistral Large 4: The Trillion-Parameter Open Flagship Arrives

Following early leaks, Mistral launched its trillion-parameter flagship Large 4 (Le Chonk) as a public research preview, with open weights due by month's end; early benchmarks praise it but question its value proposition.

2026-10-06 ~ 2026-10-11 · 7 episodes · 108 posts

Episode 1 · Mistral to Launch New Flagship Model, Claims Edge Over Chinese Models (2026-10-06, 2 posts)

Reuters reports that Mistral is set to release a new large flagship model, with CEO Arthur Mensch claiming it surpasses Chinese models in some areas, including cybersecurity, as a European achievement. Reddit commenters were largely skeptical of the claim.

Episode 2 · Mistral Launches Mistral Large 4, a 1T-Parameter Open-Weight Flagship (2026-10-06, 86 posts)

Mistral announced its flagship open-weight model Mistral Large 4 (codename Le Chonk) on October 6 as a Research Public Preview: a native multimodal MoE with 1T total parameters, 49B active, text and image input, and a 1M-token context window. The API launched the same day, with open weights expected by end of October — the Hugging Face pre-release page (relayed by victormustar) shows an expected October 31 release and lists active parameters as 52B. It is Mistral's largest open-weight flagship and a landmark move for the European open-source camp.

Confirmed

  • Specs: 1T total / 49B active parameters (HF page says 52B), native multimodal MoE, text+image input, 1M-token context; visual grounding claimed to beat some closed frontier models.
  • Positioning: the company calls it the best open-weight model on aggregate benchmarks in the US/Europe — and outside China; SOTA on cybersecurity defense, manufacturing and finance workloads, optimized for coding and cyber defense.
  • Training: co-founder Guillaume Lample disclosed that ML4 was trained end-to-end on 3,800 NVIDIA Grace Blackwell GPUs at a self-built European datacenter in Bruyères-lès-Châtel; NVIDIA called it "frontier-class compute made in Europe." Lample also teased upcoming Series C/D clusters.
  • Release cadence: API live at launch, currently Research Public Preview; weights due end of October (expected Oct 31). RL training reportedly continues to bring gains.
  • Human evals: Lample said ML4 beats GLM 5.3 on STEM, CAD and finance tasks and ties on agentic coding. LMArena's live channel Le Chonk ran a round of real-world evaluations.
  • Reception: per WIRED (relayed by 233C), Mistral positions the model as Europe's answer to Chinese open-weight rivals like DeepSeek and Qwen. Tansu Yegen argued Europe need not beat OpenAI or Google everywhere but should not be excluded from the AI race. Community response leaned memetic ("chonky"). TestingCatalog noted Google and OpenAI model updates the same day.

Unconfirmed

  • Community relays note gaps in official benchmark data; independent evaluations are pending, and detailed LMArena results have not been published.
  • Claims such as "beats GLM 5.3," "surpasses closed models," and "best outside China" come from official statements without third-party verification. The "competitor Astra used 100k GPUs" comparison cited in relays (steipete, beffjezos) comes from a third-party tweet, not official data. An earlier relayed claim of "second in blind coding tests, behind Opus 5" lacks a primary source and is not adopted.

Why it matters

  • A 1T-parameter open-weight model is a landmark for the European open-source camp, directly benchmarked against US and Chinese frontier open models.
  • Training a trillion-parameter model end-to-end on just 3,800 Grace Blackwell GPUs inside Europe — if performance claims hold up to third-party verification — challenges the assumption that frontier models require massive compute, and could reshape enterprise model selection.

66 more related posts →

Episode 3 · Mistral Large 4 Still in RL Training, Expected by Month's End (2026-10-06, 3 posts)

Mistral Large 4 (codenamed "le chonk") is still undergoing RL training with continuous improvements over its preview, and is expected to be released for wider use later this month, according to leaks and official channels.

Episode 4 · Mistral Large 4 Review: France Back in Third Place, but Value Questioned (2026-10-06, 11 posts)

On Oct 6, Artificial Analysis published its full review of Mistral Large 4 Research Public Preview: a 1T-parameter model (49B active) scoring 38 on the Intelligence Index, placing it near the Gemini/GPT tier, which the evaluator framed as France rejoining the strongest models outside the US and China. Weights are planned to open-source by month's end, but high pricing and token-efficiency concerns have sparked value-for-money controversy.

Confirmed

  • Specs and overall: 1T total parameters (49B active), Intelligence Index 38, near the Gemini/GPT tier; full results page is live
  • Country ranking: at 38, France rises to third, behind the US (Claude Opus 5.5, 58) and China (MiMo-V2.6-Pro, 46)
  • Pricing: $1.36/$4.18 per million input/output tokens, cached input $0.14; 50% off for the first two weeks brings per-task cost to $0.57, but the full-rate per-task cost is about $1.13, roughly 4x comparable open models
  • Value comparison: $1.13 per task for 38 points vs MiMo v2.6 Pro at $0.13 for 46; same score as GPT 6 Luna but 16x the cost ($0.07)
  • Cyber capability: Cyber Index 50, tied with GLM-5.3-Flash and below MiMo-V2.6-Pro (56); evaluator expects a top-three open-model cyber ranking once weights ship
  • Vulnerability test: against its 'strongest open model for cybersecurity' claim, a third-party VulnPR-100 run (100 PRs with known vulnerabilities, 1 hour each) found only 35/100, worse than Chinese open models; reshared by Mistral co-founder Guillaume Lample
  • Researcher Yuchen Jin noted it clearly trails GLM-5.3 on the AA Intelligence Index, joking about an 'eval crisis'
  • Independent evaluator haider found its output token usage per task is over 2x GPT-6 sol and Astra
  • Community reaction: BLUECOW009 mocked it as 'le chaton skinny,' questioning the value of the 1T scale

Unconfirmed

  • The month-end release of 1T-scale weights remains a stated plan only
  • Third-party opinions on the 38 score diverge: LuminaBench called it poor for a 1T model, while others note parity with GLM 5.x-tier rivals

Why it matters

  • An on-schedule open-sourcing of 1T-scale weights could reshape the open-model landscape and cement France as the third pole beyond the US and China
  • The capability-cost mismatch is the core dispute: mediocre scores at multiples of rival cost and token usage may limit real adoption
  • Cyber and other specialized capabilities still trail top Chinese open models like MiMo-V2.6-Pro, worth tracking after weights release

Episode 5 · Mistral Large 4 Teased on Hugging Face: 1T-Parameter Open-Weight Model Coming This Month (2026-10-06, 2 posts)

An upcoming-release page for Mistral Large 4, nicknamed Le Chonk, appeared on Hugging Face, teasing a natively multimodal model with 1 trillion total parameters and 52B active parameters, with open weights expected by the end of the month.

Episode 6 · Mistral Claims Large 4 Among Strongest Cybersecurity AI Models (2026-10-07, 2 posts)

Mistral AI claims its latest model Mistral Large 4 is among the world's strongest AI models for cybersecurity, though no benchmark data accompanied the claim; security firm ReasonCoreAI has launched SciCode testing to evaluate it.

Episode 7 · Mistral Large 4 Opens Public Beta: Trillion-Parameter MoE with Open Weights Coming (2026-10-10, 2 posts)

Mistral has opened a public preview of Mistral Large 4, a natively multimodal sparse MoE model with about 1 trillion total parameters and only 49 billion active, with open weights promised by the end of the month.