OpenAI quietly boosted GPT-6 Astra benchmark numbers and kept changing them after launch, Fortune reports

jeremyakahn · x · 2026-09-05

Fortune reporter Emily Forlini reports that OpenAI has revised several evaluation benchmarks for its GPT-6 Astra model since publishing the launch blog post on Sept. 3 — in some cases making Astra's numbers better while figures for rival Anthropic's models got worse.

The rollout itself was rocky: the post was slated for 2 p.m. ET but took nearly two more hours to become widely viewable, and the link OpenAI's X account tweeted at 3:32 p.m. returned errors before CEO Sam Altman posted it himself, noting "We hit a little snag."

The report raises questions about vendor-reported benchmark integrity and post-launch metric manipulation.

Original post →

More from Models

Models channel →