CMU Releases cua-speedrun: Benchmarking Speed and Cost of Computer-Use Agents

A CMU team has released cua-speedrun, an evaluation system that measures both how long a computer-use agent takes to complete tasks and how much its model calls cost, whereas most existing benchmarks only record whether tasks succeed. Authors @kohjingyu and @arankomatsuzaki note the project is supported by @scsatcmu, led by Jing Yu Koh, Pranjal Aggarwal, and Lawrence Jang, with collaborators including Sean Welleck, Daniel Fried, and Ruslan Salakhutdinov.

Confirmed

Why it matters

2026-10-01 ~ 2026-10-02 · 5 related posts

Primary sources

1 near-duplicate retellings: kohjingyu