OSWorld 2.0 Benchmark: Computer-Control Agents Remain Highly Unreliable
pstAsiatech · x · 2026-07-06
The OSWorld 2.0 benchmark evaluates computer-control agents on long-horizon real-world tasks. The paper reveals that current agents are still far from reliable in computer use: even the strongest configuration (Claude Opus 4.8 with maximum thinking and batch tool calling) achieves only a 20.6% binary accuracy and 54.8% partial score accuracy. Performance drops significantly as task length increases, with the worst performance observed when recovering hidden states, tracking massive entries, handling conflicting information, or adapting to changing requirements.
More from coding & agent
- Paper argues graph topology can become the core operating system for AI agents — theomitsa · 2026-07-27
- Claude Code desktop adds UI markup feedback for smoother visual editing — EricBuess · 2026-07-27
- Anthropic says Claude Code can drop 80% of its system prompt with no coding loss — krishnan · 2026-07-27
- Codex hit token limits during large-codebase refactors — bytebot · 2026-07-27
- Hermes agent wins praise as a browser-control harness for local models — Teknium · 2026-07-27
- A builder wants AI to reverse-engineer viral video effects into ComfyUI workflows — stale2000 · 2026-07-27