Replay agents hit SOTA on CUA benchmarks: NeurIPS oral paper exposes eval flaws
proceduralia · x · 2026-09-29
A study on computer-use agent (CUA) evaluation — selected for a NeurIPS oral (112 of 30,709 submissions) by a Meta Superintelligence Labs team — shows that a "replay agent" merely memorizing a frontier model's successful trial can reach SOTA under common protocols. The authors mathematically prove this replay agent is the agentification of pass@k, exposing the metric's flaws. They propose the PRISM principles for benchmark design and build Digiworld, a large mobile-use benchmark that auto-generates and checks millions of task variants to defeat replay attacks, plus a statistical framework for rigorous estimation. Evaluating frontier models in this environment reveals they are brittle to trivial variations (e.g., UI colors) that wouldn't matter to humans.
Related event: NeurIPS oral paper exposes replay-cheating flaw in CUA benchmarks(2 posts)→
More from Models
- ElevenLabs v4 Voice Model Now Available on Runway Platform — runwayml · 2026-09-29
- Anthropic ships Sonnet 5.5 at half the price, nearly matching Opus 5.5; Haiku 5.5 on the way — oran_ge · 2026-09-29
- Latent Space: TypeSafe CEO on why he rejects public benchmarks for Jev — Latent Space · 2026-09-29
- Xiaomi MiMo-V2.6-Distill-Qwen-9B GGUF quantization trends on Hugging Face — bartowski · 2026-09-29
- Grok 4.7 hits Amazon Bedrock: 500K context, 2x output tokens for the gains — AWS ML Blog · 2026-09-29
- OpenAI reportedly cancels October release of GPT-6.1 Astra over safety concerns — Polymarket · 2026-09-29