OSWorld 2.0 Benchmark: Computer-Control Agents Remain Highly Unreliable

pstAsiatech · x · 2026-07-06

The OSWorld 2.0 benchmark evaluates computer-control agents on long-horizon real-world tasks. The paper reveals that current agents are still far from reliable in computer use: even the strongest configuration (Claude Opus 4.8 with maximum thinking and batch tool calling) achieves only a 20.6% binary accuracy and 54.8% partial score accuracy. Performance drops significantly as task length increases, with the worst performance observed when recovering hidden states, tracking massive entries, handling conflicting information, or adapting to changing requirements.

Original post →

More from coding & agent

coding & agent channel →