OpenAI quietly changed GPT-6 Astra benchmark results around launch, Fortune reports

mark_k · x · 2026-09-05

Per Fortune, OpenAI quietly altered several GPT-6 Astra benchmark figures around launch, some changes favoring Astra and hurting rivals. Examples: Astra's hallucination rate went from 4.2% to 2%, then back to 4.2%; Anthropic Fable 5.1's FrontierMath score moved from 87.8% to 78%, then to 83%. This happened during the strange delay when Astra's launch blog was published, pulled, and republished with different numbers. OpenAI says the delay was unrelated and evals can shift with checkpoints, harnesses and reasoning configs.

Related event: OpenAI Quietly Altered GPT-6 Astra Benchmark Scores: Fortune(3 posts)→

Original post →

More from Models

Models channel →