OpenAI quietly rewrote GPT-6 Astra benchmark numbers, Fortune reports: hallucination rate 4.2% changed to 2%

mark_k · x · 2026-09-05

According to Fortune, OpenAI quietly altered several benchmark results for its GPT-6 Astra model after publishing the launch blog on Sept 3, with some changes favoring Astra and hurting rival Anthropic's numbers:

OpenAI has not publicly responded yet, and the incident raises fresh questions about the reliability of vendor-reported benchmarks.

Related event: OpenAI Quietly Altered GPT-6 Astra Benchmark Scores: Fortune(3 posts)→

Original post →

More from Models

Models channel →