Mistral Large 4 preview: 1T-parameter open MoE betting on security work others refuse

大模型之路 · wechat · 2026-10-10

Mistral launched the Mistral Large 4 public preview (nicknamed LeChonk) on Oct 6: a fine-grained MoE with 1.05T total parameters and 49B active per token, native multimodality (1.6B vision encoder), 1M-token context, and open weights planned for Oct 27 after red-teaming. It was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in European data centers, claimed to use 2-3x less compute than US/China peers.

The differentiation: filling the refusal gap. Closed frontier models largely refuse vulnerability reproduction, binary analysis, and CTF tasks (near-zero on CyberGym-E2E for Claude Opus 5.5 and GPT-6 Astra). Mistral tunes its alignment filters to distinguish defensive vulnerability fixes from malicious exploitation: 82% on open-source vuln reproduction/patching, 93% on Cybench, and 93.3% attack blocking on Lakera B3. API pricing: $0.68/M input tokens, $2.09/M output.

Weaknesses and self-hosting math. Artificial Analysis scores it 38 (up from Large 3's 9) but well below Claude Opus 5.5's 58; coding lags closed leaders by 40+ points (Terminal-Bench 22.7%). FP8 weights alone need 1.05TB memory: 8xH200 handles only 32K context, 128K needs 16xH200, and full 1M-context serving requires 16xB200. The article also flags four caveats: all numbers are vendor-reported, weight release could slip, the red-team-window version with "reduced vetting, expanded cyber capabilities" for select partners raises openness concerns, and open weights mean no provider liability.

Related event: Mistral Large 4 Opens Public Beta: Trillion-Parameter MoE with Open Weights Coming(2 posts)→

Original post →

More from Infra

Infra channel →