OpenAI quietly changed GPT-6 Astra benchmark results around launch, Fortune reports
mark_k · x · 2026-09-05
Per Fortune, OpenAI quietly altered several GPT-6 Astra benchmark figures around launch, some changes favoring Astra and hurting rivals. Examples: Astra's hallucination rate went from 4.2% to 2%, then back to 4.2%; Anthropic Fable 5.1's FrontierMath score moved from 87.8% to 78%, then to 83%. This happened during the strange delay when Astra's launch blog was published, pulled, and republished with different numbers. OpenAI says the delay was unrelated and evals can shift with checkpoints, harnesses and reasoning configs.
Related event: OpenAI Quietly Altered GPT-6 Astra Benchmark Scores: Fortune(3 posts)→
More from Models
- Your 99% Benchmark Score Is a System Score: Why GPT-6 Astra Numbers Blur Model vs Harness — algo_diver · 2026-09-06
- DiffusionGemma Hits ~6k tok/s on H200 Estimate, Making Uno Paper's Plot 'Extremely Suspicious' — bodonoghue85 · 2026-09-06
- Naval Amplifies DeepSeek Explainer: SFT Is a Bike Manual, RL Is Learning to Ride — McDonaghMatthew · 2026-09-06
- Blogger pegs 30% odds OpenAI already solved Navier-Stokes, 50% partial progress — scaling01 · 2026-09-05
- LLMs write locally coherent but globally incoherent quests: an MMO writer's thousands-of-quests problem — HLCYSWAP · 2026-09-05
- GPT-6 'Astra' at capacity? User burns 9% of weekly quota in 20 hours then hits rate limit — sick_burns2000 · 2026-09-05