Artificial Analysis Updated Benchmarks Twice in 4 Days for Astra
py-net · reddit · 2026-09-09
A Reddit post claims Artificial Analysis updated its benchmark twice in 4 days to reflect Astra's real strength. Something similar happened with Sol, after which Arena restructured its coding benchmark, added a full-stack bench, and updated the webdev bench to better reflect real-world performance. The poster concludes OpenAI seems to be the only lab not benchmaxxing.
More from Models
- OpenAI claims AI found analytical proof of Navier-Stokes blowup, verified in Lean — Dr_Singularity · 2026-09-09
- OpenAI Claims Navier-Stokes Millennium Prize Proof Produced by Agent Swarm on Next-Gen Model — daniel_mac8 · 2026-09-09
- Navier-Stokes solved in 5 days by OpenAI? Insiders say impossible things are becoming possible — MoonL88537 · 2026-09-09
- OpenAI burned millions in compute on $1M prize: 10,000 agents, 88 hours, 130B tokens — Hesamation · 2026-09-09
- OpenAI claims agent group using next-gen model solved the Navier-Stokes Millennium Prize Problem — polynoamial · 2026-09-09
- "Significantly more capable than GPT-6 Astra": OpenAI's internal model gap — kimmonismus · 2026-09-09