OpenAI vs. Anthropic Math Duel: Astra's Lead Caught in 24 Hours
新智元 · wechat · 2026-08-03
Last weekend, OpenAI and Anthropic engaged in a fierce AI duel at the frontier of mathematics.
Timeline:
- OpenAI's Move: On Aug 1, an OpenAI researcher announced that their unreleased next-gen model, Astra, successfully proved 10 long-standing open math problems, providing Lean certificates. The marginal inference cost was reportedly under $2,000.
- Anthropic's Counter: Within 24 hours, an Anthropic researcher responded that their publicly available model, Fable, independently reproduced 5 of those proofs in a clean, offline environment.
Implications:
- Shrinking Halflife: Priority in math breakthroughs used to be measured in years; now, an AI model's lead lasts mere hours.
- Cross-Verification: Two independent models arriving at the same conclusion acts as a peer-review mechanism, validating AI's mathematical reliability.
- The Verification Bottleneck: As AI rapidly generates frontier proofs, the scarce resource is shifting towards human experts capable of understanding and verifying them.
More from Models
- A user says DeepSeek cost 7 cents and beat GPT-5.6 Sol on the same task — yacineMTB · 2026-08-04
- Poster says DeepSeek outperformed GPT-5.6 Sol on this output — yacineMTB · 2026-08-04
- Kimi and GLM 5.2 pricing keeps falling as models port across hardware platforms — markjeffrey · 2026-08-04
- OpenAI Reveals How It Built Its Realtime Voice AI System in Just 6 Months — borowcy · 2026-08-04
- Frontier models still fail basic PDE solvers, benchmark post says — GaryMarcus · 2026-08-04
- DeepSeek V4 Flash beats Qwen 3.8 Max and GLM 5.2 on coding benchmarks — airesearch12 · 2026-08-04