Mistral Large 4 posts strong STEMBench score across six research domains
GuillaumeLample · x · 2026-10-07
Mistral CTO Guillaume Lample boosted a STEMBench evaluation: Mistral Large 4 scored well across 6 domains with 20 problems each, showing it can perform original research across fields with fairly even performance. The author calls it a great win for European AI.
More from Models
- Google's Gemini 3.8 Live Speech-to-Speech Models Top Voice Rankings at $0.84/Hour Input — DeepLearningAI · 2026-10-07
- Mistral's real news isn't Large 4 — it's a new architecture to ship models faster — _AndrewZhao · 2026-10-07
- Mistral Large 4 arrives: one 3D prompt benchmarked across 8 flagship models — GuillaumeLample · 2026-10-07
- Same prompt, different results: nanobanana image generation compared day over day — leslysandra · 2026-10-07
- Bug Hunt Bench: Mistral Large 4 fixes only 15/105 planted bugs, trails Qwen and Kimi — PawelHuryn · 2026-10-07
- Mistral Large 4 Fixes 15 of 105 Planted Bugs, Trails Qwen and DeepSeek in Real-Repo Test — PawelHuryn · 2026-10-07