Mistral's two-open-model agent hits 82% on ArXivLean using 10M output tokens per problem
scaling01 · x · 2026-09-12
A new ArXivLean leaderboard entry shows Mistral's agent stack—Leanstral 1.5 as prover plus Kimi K3 as coordinator—scoring a massive 82%, at the cost of roughly 10 million output tokens per problem.
The poster quips that this open-source combo outperformed GPT-6-Astra at the same token budget, and wonders what GPT-6-Astra could do with 130B tokens.
Related event: 6B Open-Source Leanstral 1.5 Beats GPT-6 Astra on ArXivLean(4 posts)→
More from Models
- Zvi: Mythos's 'Outward Statements Not Reflecting Internal State' Look Like CoT Crafted to Fool Auditors — TheZvi · 2026-09-12
- The real worry in Anthropic's report: no way to stop distillation by rivals — kimmonismus · 2026-09-12
- Meta's Confidential Computing Promise Doesn't Protect Data if Inference Isn't Attested — signulll · 2026-09-12
- Among a Flood of New AI Labs, the Humans Team Ships Its First Release — niloofar_mire · 2026-09-12
- Agent traces reveal hour-long costly 'super resolution' rabbit hole in continual learning study — yuxiangw_cs · 2026-09-12
- Gary Marcus mocks ChatGPT regression: is this "AGI"? — GaryMarcus · 2026-09-12