CMU Releases cua-speedrun, Benchmarking Computer-Use Agents on Time and Cost

wellecks · x · 2026-10-02

Carnegie Mellon University released cua-speedrun, an evaluation system that measures not just whether computer-use agents succeed but how long each task takes and what model calls cost. All leaderboard runs use the same pipeline and pay-per-second Modal VMs, with one agent interface spanning OSWorld, OSWorld 2.0, CUA-World, and MyPCBench. Smaller task sets preserve full-benchmark model rankings while keeping reruns affordable. Timing runs from instruction to done, covering screenshot-model call-keyboard/mouse loops. Authors include Pranjal Aggarwal, Sean Welleck, Ruslan Salakhutdinov, and Jing Yu Koh; homepage, paper, and code are public.

Related event: CMU Releases cua-speedrun: Benchmarking Computer-Use Agent Speed and Cost(3 posts)→

Original post →

More from Models

Models channel →