OpenAI vs. Anthropic Math Duel: Astra's Lead Caught in 24 Hours
新智元 · wechat · 2026-08-03
Last weekend, OpenAI and Anthropic engaged in a fierce AI duel at the frontier of mathematics.
Timeline:
- OpenAI's Move: On Aug 1, an OpenAI researcher announced that their unreleased next-gen model, Astra, successfully proved 10 long-standing open math problems, providing Lean certificates. The marginal inference cost was reportedly under $2,000.
- Anthropic's Counter: Within 24 hours, an Anthropic researcher responded that their publicly available model, Fable, independently reproduced 5 of those proofs in a clean, offline environment.
Implications:
- Shrinking Halflife: Priority in math breakthroughs used to be measured in years; now, an AI model's lead lasts mere hours.
- Cross-Verification: Two independent models arriving at the same conclusion acts as a peer-review mechanism, validating AI's mathematical reliability.
- The Verification Bottleneck: As AI rapidly generates frontier proofs, the scarce resource is shifting towards human experts capable of understanding and verifying them.
More from Models
- ChatGPT co-inventor launches Jev: claims 20-200x faster, 40-400x cheaper than LLMs — GabGarrett · 2026-09-18
- Ran Jev across 10 services for hours, still couldn't spend $1 — multiply_matrix · 2026-09-18
- Tencent's Hy4 Preview ranks #4 among open-weight models, cheapest in top ten — mariofilhoml · 2026-09-18
- Simple letter-counting test exposes huge gap: GPT-6-Astra hits 93%, Fable 5.1 flounders — scaling01 · 2026-09-18
- Jev, a 'System One' model by Typesafe, launches on OpenRouter with typed decisions instead of text — majidmanzarpour · 2026-09-18
- Zhipu claims AI autonomously discovered a WeWorm-exploitable vulnerability — teortaxesTex · 2026-09-18