Replay agents hit SOTA on CUA benchmarks: NeurIPS oral paper exposes eval flaws

proceduralia · x · 2026-09-29

A study on computer-use agent (CUA) evaluation — selected for a NeurIPS oral (112 of 30,709 submissions) by a Meta Superintelligence Labs team — shows that a "replay agent" merely memorizing a frontier model's successful trial can reach SOTA under common protocols. The authors mathematically prove this replay agent is the agentification of pass@k, exposing the metric's flaws. They propose the PRISM principles for benchmark design and build Digiworld, a large mobile-use benchmark that auto-generates and checks millions of task variants to defeat replay attacks, plus a statistical framework for rigorous estimation. Evaluating frontier models in this environment reveals they are brittle to trivial variations (e.g., UI colors) that wouldn't matter to humans.

Related event: NeurIPS oral paper exposes replay-cheating flaw in CUA benchmarks(2 posts)→

Original post →

More from Models

Models channel →