Mistral Large 4 ties #1 on 669 clinical decisions with zero severe misses
GuillaumeLample · x · 2026-10-07
Mistral Large 4 (aka Le Chonk, 1T params, 49B active, natively multimodal) scored 617/669 correct on clinical decisions on day one, tying for #1 with Jev, Clef and pplx-decider. Notably, it is the only top model with zero severe misses — no missed emergencies and no allergy history misread as medication taken. On triage it scored 238/240, the best so far. Available via API now; open weights release end of October.
Related event: 24 Models Benchmarked on 669 Clinical Decision Tasks(3 posts)→
More from Models
- Grok 4.7 goes live on Microsoft Foundry, expanding xAI's enterprise reach — SpaceXAI · 2026-10-07
- Claude turns OpenAI's 53-page Hodge conjecture proof into a 3D animation — imjustnewatai · 2026-10-07
- Claude Opus 5.5 tops AA Index at 58, but per-task cost stays flat as output volume jumps 60% — qinzytech · 2026-10-07
- Small LLMs Hit Only 15% Rank-Score Consistency in Financial Analysis Tests vs 75% for Frontier Models — Rough_Practice7631 · 2026-10-07
- Cited by AGI ranks human mathematicians most cited across OpenAI's /math manuscript collection — willdepue · 2026-10-07
- Seroter's reading list: Mistral's 1T-param Large 4 and a context quality framework — rseroter · 2026-10-07