OpenAI quietly rewrote GPT-6 Astra benchmark numbers, Fortune reports: hallucination rate 4.2% changed to 2%
mark_k · x · 2026-09-05
According to Fortune, OpenAI quietly altered several benchmark results for its GPT-6 Astra model after publishing the launch blog on Sept 3, with some changes favoring Astra and hurting rival Anthropic's numbers:
- Astra's reported hallucination rate was changed from 4.2% to 2%, then back to 4.2%
- Anthropic Fable 5.1's FrontierMath score moved from 87.8% to 78%, then shifted again
- The rollout itself was unusual: the blog post was slated for 2 p.m. ET but only became widely viewable nearly two hours later
OpenAI has not publicly responded yet, and the incident raises fresh questions about the reliability of vendor-reported benchmarks.
Related event: OpenAI Quietly Altered GPT-6 Astra Benchmark Scores: Fortune(3 posts)→
More from Models
- Astra Benchmark Performance Still Poor, Says Early Tester — gleech · 2026-09-06
- LLMs Play Chess Poorly, So This One Built Its Own Chess Engine to Fight Back — MikePFrank · 2026-09-06
- Astra Computer Access Fails Drag-and-Drop Web Page Build Test — BLUECOW009 · 2026-09-06
- Astra disappoints on harness-building tasks while Fable 5.1 excels, dev reports — HarveenChadha · 2026-09-06
- GPT-6 Astra runs DOOM at 20+ FPS on its own CPU, up from GPT-5.6 Sol's 1 FPS — Angaisb_ · 2026-09-06
- GPT-6 Astra designs a working jet plant and ships a live 3D simulation autonomously — deanwball · 2026-09-06