1MB replay script beats frontier models: NeurIPS oral paper exposes flaws in CUA benchmarks

proceduralia · x · 2026-09-29

A paper on evaluating computer use agents (CUAs) was selected for a NeurIPS oral presentation — only 112 orals out of 30,709 submissions.

Key finding: a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks; the authors prove its expected success rate exactly equals the source agent's pass@k in deterministic environments.

Root causes: non-principled environment design (static, unsandboxed, unreliably verified) and flawed evaluation methodology (naive aggregation, misuse of pass@k for stateful UI interactions).

Contributions:

Related event: NeurIPS oral paper exposes replay-cheating flaw in CUA benchmarks(2 posts)→

Original post →

More from Research

Research channel →