New OEIS Open Benchmark: Opus 4.8 Solves 30% of Unsolved Math Conjectures
xeophon · x · 2026-08-12
tmkadamcz released OEIS Open, a new mathematical benchmark comprising 492 unsolved math conjectures formalized in the Lean language.
Benchmark results show that, when provided with simple tools and a budget of $50 per conjecture, Claude Opus 4.8 can successfully generate Lean proofs to resolve 30% of the problems. This marks a significant breakthrough for large language models in advanced mathematical reasoning and formal theorem proving.
More from Models
- LiquidAI Launches 3B Vision-Language Model LFM2.5-VL, Outscoring Larger Rivals — JosephJacks_ · 2026-08-13
- 21-Year-Old Math Enigma Solved by Human; GPT and Claude Both Failed — anshulkundaje · 2026-08-13
- Reviewing AI Like an Art Critic: Grok 4.6 Tested on Astrology & Philosophy — karinanguyen · 2026-08-13
- Kimi K3 Hits GitHub Copilot: Developers Build Games from a Single Prompt — film_girl · 2026-08-13
- MiniMax H3 vs SD2.5: H3 Delivers Smoother Motion and Better Commercial Visuals — FellMentKE · 2026-08-13
- Bittensor Subnet Runs Full Kimi K3 on 80 RTX 5090s, Halving API Costs — markjeffrey · 2026-08-13