CMU ships cua-speedrun: standardized benchmark finally measures computer-use agent speed and cost
rsalakhu · x · 2026-10-05
CMU researchers (Lawrence Jang, Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried, et al.) argue computer-use agent evaluation is a mess with a reproducibility crisis: task success is comparable across setups, but runtime depends heavily on the eval infrastructure, making progress on faster CUAs unmeasurable. Their fix, cua-speedrun, is a standardized system that runs evaluations like OSWorld/OSWorld 2.0 on identical Modal-hosted VMs with one agent interface, recording not just success but per-task time and model-call cost. A curated 50-task subset keeps rankings aligned with full benchmarks while making reruns cheap; the site demos eight agents racing through an average OSWorld task at 40x speed, ranked by completion time.
More from coding & agent
- Open Instinct: MIT-licensed personal-agent policy engine with allow/ask/deny outcomes — maritime_sh · 2026-10-05
- Skills vs MCP vs RAG vs Memory: a 4-part framework for agent knowledge — MaryamMiradi · 2026-10-05
- Model-written tests rejected a known-correct solution 77% of the time in agent pipeline test — deadatreides1 · 2026-10-05
- GitHub Copilot CLI v1.0.92-4 adds config subcommands and a batch of stability fixes — copilot-cli-release-app[bot] · 2026-10-05
- Sourcegraph CEO: AI-generated code is decaying large codebases at scale — AI Engineer · 2026-10-05
- Inner local agent injects context into edge agents via attestation — natesiggard · 2026-10-05