Agent Success Rate Drops on Repeat: Paper Reveals Computer Use Reliability Trap

xwang_lk · x · 2026-08-12

Current industry benchmarks like OSWorld typically report a single success rate (e.g., 85%), which fails to reflect the actual requirements for enterprise deployment.

Researchers tested identical computer-use tasks 10 times each. Results show that pass@10 (succeeding at least once in ten tries) reached 78%, but pass^10 (succeeding on all ten consecutive tries) plummeted to 36%. This highlights a massive gap between model capability and deployment reliability, as real-world business processes rarely allow for multiple retries.

The paper On the Reliability of Computer Use Agents analyzes three main factors contributing to agent unreliability: execution stochasticity, task specification ambiguity, and behavioral variability. It suggests evaluating agents under repeated executions and favoring stable strategies like neuro-symbolic agents.

Original post →

More from coding & agent

coding & agent channel →