CMU Releases cua-speedrun, Benchmarking Computer-Use Agents on Time and Cost
wellecks · x · 2026-10-02
Carnegie Mellon University released cua-speedrun, an evaluation system that measures not just whether computer-use agents succeed but how long each task takes and what model calls cost. All leaderboard runs use the same pipeline and pay-per-second Modal VMs, with one agent interface spanning OSWorld, OSWorld 2.0, CUA-World, and MyPCBench. Smaller task sets preserve full-benchmark model rankings while keeping reruns affordable. Timing runs from instruction to done, covering screenshot-model call-keyboard/mouse loops. Authors include Pranjal Aggarwal, Sean Welleck, Ruslan Salakhutdinov, and Jing Yu Koh; homepage, paper, and code are public.
Related event: CMU Releases cua-speedrun: Benchmarking Computer-Use Agent Speed and Cost(3 posts)→
More from Models
- Animation Bench: GPT-6.1 Sol tops coding agents on web animation reconstruction at $0.46/task — himanshustwts · 2026-10-02
- GPT-6.1 Sol tops Animation Bench, first model to cross 0.5 on motion consistency — himanshustwts · 2026-10-02
- Meta's Muse Stuns Users, Helps Lift Stock 10% in a Week — alexandr_wang · 2026-10-02
- Fable 5.5 Rumored Next as Anthropic Speeds Up 0.4-Jump Version Cadence — ChrisGPT · 2026-10-02
- VAmoS Pro voice-agent benchmark: Grok leads tasks, GPT-Live fastest, Gemini most noise-robust — davlanade · 2026-10-02
- Opus 5.5 keeps saying "himbo" — a verbal quirk no previous Claude showed — repligate · 2026-10-02