cua-speedrun: CMU benchmark shows 4.4x speed gap between equal-scoring computer-use agents
arankomatsuzaki · x · 2026-10-01
CMU researchers (including Ruslan Salakhutdinov and Jing Yu Koh) released cua-speedrun, an evaluation system that measures not just whether computer-use agents succeed but how long each task takes and what model calls cost — shared widely by Andrej Karpatzumaki Komatsuzaki. All runs use the same pipeline and per-second-billed Modal VMs across OSWorld, OSWorld 2.0, CUA-World and MyPCBench, with a 50-task subset whose rankings match the full benchmarks. Early results: GPT-6 Astra averages 1:30 per task at 89.6% score while Kimi K3 needs 7:23 for the same score — a 4.4x gap, with Gemini 3.8 Flash and Claude Opus/Sonnet 5 variants also charted. The takeaway: speed and cost differences between agents are as decisive as quality.
More from coding & agent
- Priors: An Onchain Credit Bureau for AI Agents Emerges — econoar · 2026-10-01
- Tripo API integrates with Unbound for real-time sculpting of AI-generated 3D assets — yshan2u · 2026-10-01
- Veteran game dev lists 10 ways AI saves time: crash logs, CMake, 400ms frame hitches — draginol · 2026-10-01
- NVIDIA's Mid-Harness Scales Actions at the Model-Harness Boundary, Lifting TerminalBench Pass@1 to 68.03% — nvidia · 2026-10-01
- Amazon's SMART Self-Evolving Multi-Agent System Tops All 15 Subtitle Arena Directions, Cuts Penalty 6.9% — amazon · 2026-10-01
- Gary Bernhardt hits all-time low faith in AI agents: they "fix" tests by deleting them — sidjustice_ · 2026-10-01