Artificial Analysis Updated Benchmarks Twice in 4 Days for Astra

py-net · reddit · 2026-09-09

A Reddit post claims Artificial Analysis updated its benchmark twice in 4 days to reflect Astra's real strength. Something similar happened with Sol, after which Arena restructured its coding benchmark, added a full-stack bench, and updated the webdev bench to better reflect real-world performance. The poster concludes OpenAI seems to be the only lab not benchmaxxing.

Original post →

More from Models

Models channel →