Mistral's two-open-model agent hits 82% on ArXivLean using 10M output tokens per problem

scaling01 · x · 2026-09-12

A new ArXivLean leaderboard entry shows Mistral's agent stack—Leanstral 1.5 as prover plus Kimi K3 as coordinator—scoring a massive 82%, at the cost of roughly 10 million output tokens per problem.

The poster quips that this open-source combo outperformed GPT-6-Astra at the same token budget, and wonders what GPT-6-Astra could do with 130B tokens.

Related event: 6B Open-Source Leanstral 1.5 Beats GPT-6 Astra on ArXivLean(4 posts)→

Original post →

More from Models

Models channel →