Mistral Prover hits 82.61% verified solutions on MathArena's ArXivLean benchmark
AlbertQJiang · x · 2026-09-11
MathArena added Mistral Prover (Leanstral 1.5 + light K3 orchestration), scoring 38/46 verified solutions (82.61%) on ArXivLean June 2026, beating GPT-6 Astra. AlbertQJiang notes the score comes mostly from heavy grinding by the 6B Leanstral 1.5, and suspects token counting on the leaderboard is erroneous. Work by intern Matéo Pirio Rossignol, mentored by Roman Soletskyi. MathArena also added Fable 5.1 (max) and Qwen3.8-Max on Sept 7.
Related event: 6B Open-Source Leanstral 1.5 Beats GPT-6 Astra on ArXivLean(4 posts)→
More from Models
- Startup reportedly builds autonomous drone system using GPT-6 Astra to track people from a single image — Polymarket · 2026-09-12
- Nex-N2.5 Pro, a 397B multimodal model focused on Computer Use, quietly lands on OpenRouter — nikola_mr64990 · 2026-09-12
- GPT-6 fails to improve on molecular property prediction, fueling AGI skepticism — GaryMarcus · 2026-09-12
- User complains Grok Bot keeps getting dumber with use — billyjhowell · 2026-09-12
- Every Builds Internal Platform for Personal Benchmarks Based on Real Work, Not Leaderboards — danshipper · 2026-09-12
- Gemini live search repeatedly failing on basic queries, users report — Dimensional-Misfit · 2026-09-12