OSWorld 2.0 Benchmark: Computer-Control Agents Remain Highly Unreliable
pstAsiatech · x · 2026-07-06
The OSWorld 2.0 benchmark evaluates computer-control agents on long-horizon real-world tasks. The paper reveals that current agents are still far from reliable in computer use: even the strongest configuration (Claude Opus 4.8 with maximum thinking and batch tool calling) achieves only a 20.6% binary accuracy and 54.8% partial score accuracy. Performance drops significantly as task length increases, with the worst performance observed when recovering hidden states, tracking massive entries, handling conflicting information, or adapting to changing requirements.
More from coding & agent
- Goal-driven AI needs verifiable success signals, or it invents its own — daniel_mac8 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- Anthropic researcher: 99% of engineers now run swarms of 300+ self-improving agents — AlishaOutridge · 2026-09-11
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11