cua-speedrun: CMU benchmark shows 4.4x speed gap between equal-scoring computer-use agents

arankomatsuzaki · x · 2026-10-01

CMU researchers (including Ruslan Salakhutdinov and Jing Yu Koh) released cua-speedrun, an evaluation system that measures not just whether computer-use agents succeed but how long each task takes and what model calls cost — shared widely by Andrej Karpatzumaki Komatsuzaki. All runs use the same pipeline and per-second-billed Modal VMs across OSWorld, OSWorld 2.0, CUA-World and MyPCBench, with a 50-task subset whose rankings match the full benchmarks. Early results: GPT-6 Astra averages 1:30 per task at 89.6% score while Kimi K3 needs 7:23 for the same score — a 4.4x gap, with Gemini 3.8 Flash and Claude Opus/Sonnet 5 variants also charted. The takeaway: speed and cost differences between agents are as decisive as quality.

Original post →

More from coding & agent

coding & agent channel →