Astra becomes first public model to solve any curated hard Erdős problems with Lean proofs
Jsevillamol · x · 2026-09-04
Greg Burnham proposes a new AI math benchmarking idea: curate a set of roughly Erdős-hard problems, require Lean formal proofs, and run all of them with a substantial compute budget. Astra turns out to be the first public model scoring above 0% — officially 2/68 solved, plus 3 more in ad hoc runs. Using Lean proofs as an anti-cheating mechanism could shape the next generation of hard math evals.
More from Models
- Early Access Tester: GPT-6 Astra Built a Blender Werewolf in 8 Minutes and a World Simulator in 17 — TheMoonMidas · 2026-09-04
- swyx Burned 20B Tokens Stress-Testing Astra on Real AI Engineering Tasks — All for Under $6/Hour — charliermarsh · 2026-09-04
- ARC-AGI's Kamradt: clever harnesses measure human intelligence, not models — GregKamradt · 2026-09-04
- mitsuhiko on Astra: model capability gains show no sign of slowing down — mitsuhiko · 2026-09-04
- Astra is 2.5x pricier per token yet cheaper and faster per task than sol and fable — 12exyz · 2026-09-04
- How OpenAI counts 'messages' for GPT-6 Astra usage limits puzzles users — Original-League-6094 · 2026-09-04