CMU Releases cua-speedrun: Benchmarking Speed and Cost of Computer-Use Agents
A CMU team has released cua-speedrun, an evaluation system that measures both how long a computer-use agent takes to complete tasks and how much its model calls cost, whereas most existing benchmarks only record whether tasks succeed. Authors @kohjingyu and @arankomatsuzaki note the project is supported by @scsatcmu, led by Jing Yu Koh, Pranjal Aggarwal, and Lawrence Jang, with collaborators including Sean Welleck, Daniel Fried, and Ruslan Salakhutdinov.
Confirmed
- Core finding: at equal scores, computer-use agents can differ in speed by up to 4.4x.
- Additional data from @kohjingyu: different models and reasoning effort levels trade off performance, speed, and cost differently; on OSWorld-Verified, the time-performance Pareto frontier is not dominated by a single family but split among multiple models, including the Claude Opus series.
- @wellecks shared engineering details: all leaderboard runs use the same pipeline and are billed per second on Modal, ensuring time and cost measurements are uniform and comparable.
- The paper, code, and interactive leaderboard are all open source.
Why it matters
- Traditional benchmarks only look at success rates, ignoring latency and cost, which are equally critical in real deployments; cua-speedrun brings both into a unified measurement framework.
- The finding that the Pareto frontier is split across multiple models shows no single model leads on speed, performance, and cost simultaneously, providing a quantitative basis for model selection and further optimization.
2026-10-01 ~ 2026-10-02 · 5 related posts
Primary sources
- cua-speedrun: CMU benchmark shows 4.4x speed gap between equal-scoring computer-use agents — arankomatsuzaki · 2026-10-01
- [source] cua-speedrun: the time-performance Pareto frontier is split across model families — kohjingyu · 2026-10-01
- cua-speedrun goes open: leaderboard, paper, and code from CMU — kohjingyu · 2026-10-01
- [source] CMU Releases cua-speedrun, Benchmarking Computer-Use Agents on Time and Cost — wellecks · 2026-10-02
1 near-duplicate retellings: kohjingyu