GPT-6 Astra reportedly scores 3% on FrontierMath Erdős, spending $220K in compute
haider1 · x · 2026-09-05
Unconfirmed report: GPT-6 "Astra" scored 3% on FrontierMath Erdős while every other tested model (Fable 5.1, Fable 5, 5.6 Sol, 5.5) scored 0%. Across repeated attempts Astra eventually solved 5 of 68 open problems, but at a compute cost of over $220,000.
More from Models
- Yoav Goldberg: capabilities once dependent on the harness are now baked into the model — yoavgo · 2026-09-05
- Yoav Goldberg: Ark's harness was simply bad, and OpenAI's fix was obvious — yoavgo · 2026-09-05
- Anthropic Fable 5.1 vs OpenAI Astra: analyst teases a clear winner — dylan522p · 2026-09-05
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- Meta ships Muse Spark 1.3 with max reasoning, pitching frontier performance at non-frontier prices — AIatMeta · 2026-09-05
- LLMs are now making up words that don't exist, not just jargon — StewartalsopIII · 2026-09-05