Claude Opus 5 Hits 70.6% on OSWorld 2.0, Accelerating Agent Eval Catch-Up

taoyds · x · 2026-07-25

With the release of Claude Opus 5, the model achieved a high score of 70.6% on the OSWorld 2.0 benchmark for computer-control agents. Developers noted that every time a harder eval is built, models catch up faster than expected, making it increasingly difficult to keep evaluations ahead of model capabilities.

Related event: Claude Opus 5 Tops OSWorld v2 Benchmark(2 posts)→

Original post →

More from coding & agent

coding & agent channel →