CMU researchers launch cua-speedrun, a benchmark timing and costing computer-use agents

kohjingyu · x · 2026-10-01

A Carnegie Mellon team (Jing Yu Koh, Pranjal Aggarwal, Lawrence Jang, with Welleck, Fried, Salakhutdinov) released cua-speedrun, an evaluation system that measures computer-use agents on accuracy, speed, and cost — dimensions most benchmarks ignore. All runs use an identical pay-per-second Modal VM pipeline across OSWorld, OSWorld 2.0, CUA-World and MyPCBench. On a 50-task set, GPT-6 Astra (medium) finished fastest at 1:30 (89.6%), while Claude Opus 5.5 (xhigh) topped accuracy at 97.6% but took 2:32. The time-performance Pareto frontier is spread across multiple model families and reasoning-effort settings, including Opus 5.5, GPT-6 Astra, and GPT-5.6 Luna.

Related event: CMU Releases cua-speedrun: Benchmarking Speed and Cost of Computer-Use Agents(5 posts)→

Original post →

More from coding & agent

coding & agent channel →