Claude Opus 5 Hits 70.6% on OSWorld 2.0, Accelerating Agent Eval Catch-Up
taoyds · x · 2026-07-25
With the release of Claude Opus 5, the model achieved a high score of 70.6% on the OSWorld 2.0 benchmark for computer-control agents. Developers noted that every time a harder eval is built, models catch up faster than expected, making it increasingly difficult to keep evaluations ahead of model capabilities.
Related event: Claude Opus 5 Tops OSWorld v2 Benchmark(2 posts)→
More from coding & agent
- GitHub shows whatbroke, a tiny offline diff tool for AI agent behavior — JeffLadish · 2026-07-25
- A DevOps builder wants to sell AI agents a trusted sandbox for $0.06 an hour — Curious_Coder098 · 2026-07-25
- Whatbroke diffs AI agent traces to catch tool-call regressions after model swaps — Impossible-Alarm-738 · 2026-07-25
- Geist's Badges adds automatic spacing for circular icons and better contrast — evilrabbit_ · 2026-07-25
- Composable agents now support MCP with per-user dynamic authentication — irvinebroque · 2026-07-25
- Wayfinder says trading agents need a secure execution layer, not just Claude Code — templecrash · 2026-07-25