1MB replay script beats frontier models: NeurIPS oral paper exposes flaws in CUA benchmarks
proceduralia · x · 2026-09-29
A paper on evaluating computer use agents (CUAs) was selected for a NeurIPS oral presentation — only 112 orals out of 30,709 submissions.
Key finding: a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks; the authors prove its expected success rate exactly equals the source agent's pass@k in deterministic environments.
Root causes: non-principled environment design (static, unsandboxed, unreliably verified) and flawed evaluation methodology (naive aggregation, misuse of pass@k for stateful UI interactions).
Contributions:
- PRISM, five design principles for CUA environments (privileged verification, realistic environments, integrity-checked configs, sandboxed execution, multifactorial variability), instantiated in DigiWorld — a benchmark of 15 realistic sandboxed mobile apps spanning over 3.2 million verified unique configurations
- An aggregation framework pairing Wilson score intervals with hierarchical bootstrap, producing confidence intervals that correctly account for the nested structure of CUA benchmarks
Related event: NeurIPS oral paper exposes replay-cheating flaw in CUA benchmarks(2 posts)→
More from Research
- Paper2Agent turns research papers into interactive AI agents, Nature paper shows — james_y_zou · 2026-09-29
- D-JEPA proposes decision-aligned latent world model, hits 87.89% on PushT — NEBULIS-Lab · 2026-09-29
- NanoGPT speedrun record falls to 39.9s, -46% via flop-level skipping tricks — yacinelearning · 2026-09-29
- 2004 Paper Shows Leaf Stomata Perform Distributed Computation, Like Cellular Automata — eigenron · 2026-09-29
- Dev experiments: making 3D text visualization useful beyond a novelty — SnooPeripherals5313 · 2026-09-29
- giffmana calls CPC the GOAT paper: one technique shown across 4 domains, each with a follow-up — giffmana · 2026-09-29