CMU ships cua-speedrun: standardized benchmark finally measures computer-use agent speed and cost

rsalakhu · x · 2026-10-05

CMU researchers (Lawrence Jang, Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried, et al.) argue computer-use agent evaluation is a mess with a reproducibility crisis: task success is comparable across setups, but runtime depends heavily on the eval infrastructure, making progress on faster CUAs unmeasurable. Their fix, cua-speedrun, is a standardized system that runs evaluations like OSWorld/OSWorld 2.0 on identical Modal-hosted VMs with one agent interface, recording not just success but per-task time and model-call cost. A curated 50-task subset keeps rankings aligned with full benchmarks while making reruns cheap; the site demos eight agents racing through an average OSWorld task at 40x speed, ranked by completion time.

Original post →

More from coding & agent

coding & agent channel →