NeurIPS oral paper exposes replay-cheating flaw in CUA benchmarks
A NeurIPS oral paper (one of 112 among 30,709 submissions) from Meta Superintelligence Lab reveals that a mere 1MB replay script can beat frontier models on computer-use agent (CUA) benchmarks, exposing a fundamental flaw in current CUA evaluations.
2026-09-29 ~ 2026-09-29 · 2 related posts
- Replay agents hit SOTA on CUA benchmarks: NeurIPS oral paper exposes eval flaws — proceduralia · 2026-09-29
- 1MB replay script beats frontier models: NeurIPS oral paper exposes flaws in CUA benchmarks — proceduralia · 2026-09-29