Mistral Large 4 preview: 1T-parameter open MoE betting on security work others refuse
大模型之路 · wechat · 2026-10-10
Mistral launched the Mistral Large 4 public preview (nicknamed LeChonk) on Oct 6: a fine-grained MoE with 1.05T total parameters and 49B active per token, native multimodality (1.6B vision encoder), 1M-token context, and open weights planned for Oct 27 after red-teaming. It was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in European data centers, claimed to use 2-3x less compute than US/China peers.
The differentiation: filling the refusal gap. Closed frontier models largely refuse vulnerability reproduction, binary analysis, and CTF tasks (near-zero on CyberGym-E2E for Claude Opus 5.5 and GPT-6 Astra). Mistral tunes its alignment filters to distinguish defensive vulnerability fixes from malicious exploitation: 82% on open-source vuln reproduction/patching, 93% on Cybench, and 93.3% attack blocking on Lakera B3. API pricing: $0.68/M input tokens, $2.09/M output.
Weaknesses and self-hosting math. Artificial Analysis scores it 38 (up from Large 3's 9) but well below Claude Opus 5.5's 58; coding lags closed leaders by 40+ points (Terminal-Bench 22.7%). FP8 weights alone need 1.05TB memory: 8xH200 handles only 32K context, 128K needs 16xH200, and full 1M-context serving requires 16xB200. The article also flags four caveats: all numbers are vendor-reported, weight release could slip, the red-team-window version with "reduced vetting, expanded cyber capabilities" for select partners raises openness concerns, and open weights mean no provider liability.
More from Infra
- Let's Encrypt moves to 64-day certificate lifetimes by default in Feb 2027 — jedisct1 · 2026-10-11
- Anthropic engineer hits 1 billion tokens per day for a week — AaronBergman18 · 2026-10-11
- NVIDIA inference expert on when to keep, repurpose or replace aging GPUs — kimmonismus · 2026-10-11
- Independent Researcher Makes TPU Pallas top-k Bitwise Correct and 1.67x Faster — Francis_YAO_ · 2026-10-11
- AI agent tunes Triton kernels on AMD MI210, flipping grid order yields 1.29x speedup — zmkzmkz · 2026-10-11
- Engineer with $120k of GPUs: bought 75% before the price surge, don't follow me — TheZachMueller · 2026-10-11