LeanLean benchmark: Opus 5.5 scores 64.3% compressing Lean proofs, GPT 6.1 Sol only 39.9%

ChrSzegedy · x · 2026-10-07

Researchers released LeanLean, a benchmark for compressing Lean codebases, motivated by LLM-written proofs ballooning in size—Claude's Lean proof of Fermat's Last Theorem runs 13M lines.

Leaderboard: Opus 5.5 dominates with 64.3%, while GPT 6.1 Sol reaches only 39.9%.

Original post →

More from Research

Research channel →