LeanLean benchmark: Opus 5.5 scores 64.3% compressing Lean proofs, GPT 6.1 Sol only 39.9%
ChrSzegedy · x · 2026-10-07
Researchers released LeanLean, a benchmark for compressing Lean codebases, motivated by LLM-written proofs ballooning in size—Claude's Lean proof of Fermat's Last Theorem runs 13M lines.
Leaderboard: Opus 5.5 dominates with 64.3%, while GPT 6.1 Sol reaches only 39.9%.
More from Research
- Gamow Labs CEO uses AI to reanalyze rare-disease genomes, found what prenatal sequencing missed — danielmckinn0n · 2026-10-07
- Mathematician admits OAI's Hodge conjecture progress outpaced expectations, mocks the hype — ctjlewis · 2026-10-07
- Anders Sandberg: preprints are often more honest than final publications — anderssandberg · 2026-10-07
- 300B tokens on Hadwiger–Nelson: five colors ruled out, lower bound moves to 6 or 7 — soumitrashukla9 · 2026-10-07
- Two counterexamples to the Shafarevich conjecture spotted, littmath jokes about double-dipping — littmath · 2026-10-07
- Frozen 7B model with external memory cartridges hits 118/128 recalls in solo prototype — Nearby_Indication474 · 2026-10-07