OpenAI quietly boosted GPT-6 Astra benchmark numbers and kept changing them after launch, Fortune reports
jeremyakahn · x · 2026-09-05
Fortune reporter Emily Forlini reports that OpenAI has revised several evaluation benchmarks for its GPT-6 Astra model since publishing the launch blog post on Sept. 3 — in some cases making Astra's numbers better while figures for rival Anthropic's models got worse.
The rollout itself was rocky: the post was slated for 2 p.m. ET but took nearly two more hours to become widely viewable, and the link OpenAI's X account tweeted at 3:32 p.m. returned errors before CEO Sam Altman posted it himself, noting "We hit a little snag."
The report raises questions about vendor-reported benchmark integrity and post-launch metric manipulation.
More from Models
- GPT-6 Astra flunks complex PCB routing after 2h20m and 15% of weekly limits in biggest public test — yacineMTB · 2026-09-05
- Fable 5.1 medium effort matches Fable 5 high, no longer breaks prompt cache — lydiahallie · 2026-09-05
- GPT-6 Astra's first task uncovers two bugs from GPT5.6 Sol fix — op7418 · 2026-09-05
- OpenAI's 'yapping penalty' cuts Astra's HealthBench lead nearly in half after length adjustment — imjustnewatai · 2026-09-05
- GPT-6 Astra demoed redesigning board components and traces — jasonkneen · 2026-09-05
- Investor: GPT-6 Astra shows OpenAI and Anthropic are far ahead, Gemini 'benchmaxxed' — firstadopter · 2026-09-05