Agent Success Rate Drops on Repeat: Paper Reveals Computer Use Reliability Trap
xwang_lk · x · 2026-08-12
Current industry benchmarks like OSWorld typically report a single success rate (e.g., 85%), which fails to reflect the actual requirements for enterprise deployment.
Researchers tested identical computer-use tasks 10 times each. Results show that pass@10 (succeeding at least once in ten tries) reached 78%, but pass^10 (succeeding on all ten consecutive tries) plummeted to 36%. This highlights a massive gap between model capability and deployment reliability, as real-world business processes rarely allow for multiple retries.
The paper On the Reliability of Computer Use Agents analyzes three main factors contributing to agent unreliability: execution stochasticity, task specification ambiguity, and behavioral variability. It suggests evaluating agents under repeated executions and favoring stable strategies like neuro-symbolic agents.
More from coding & agent
- xAI Launches Grok Bot: Autonomous AI Agents for Real-World Workflows — XFreeze · 2026-08-12
- KohakuTerrarium: Batteries-Included Framework for Multi-Agent Teams — tom_doerr · 2026-08-12
- Build a 3D Brand Logo Material Playground Instantly with Lovable — felixhhaas · 2026-08-12
- Gemini for Go Developers: Model Selection and Agent Development Guide — rseroter · 2026-08-12
- Mojo 1.0 Released: The Systems Language for the AI Era — clattner_llvm · 2026-08-12
- AI Agent Breaks Out of 'Air Force One' Level Sandbox to Book a Flight — sloppenheimer · 2026-08-12