CMU's cua-speedrun finds Jev 10x less accurate and 1.5x slower than Opus-5.5 for computer-use agents
wellecks · x · 2026-10-03
CMU researchers launched cua-speedrun, an evaluation system that measures not just whether computer-use agents (CUA) succeed, but how long each task takes and what model calls cost — and their Jev results are surprising:
- 1.5x slower than Opus-5.5-xHigh
- 10x lower accuracy
- 5x more steps
- Overall in GPT-4 league of performance, Astra-xHigh on time
How cua-speedrun works:
- All leaderboard runs use the same pipeline and VM setup on Modal (pay-per-second cloud), with a single agent interface across OSWorld, OSWorld 2.0, CUA-World, and MyPCBench
- Smaller curated task sets keep rankings consistent with full benchmarks while making repeated runs affordable
- Timing covers the full loop: screenshot, model call, keyboard/mouse actions, with hardware, desktops, tasks, and the clock held fixed so differences are attributable to the agent
Team: Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh (CMU).
More from Models
- Claude Opus 5.5 Max tops WebDev Arena with 838K votes across 138 models — arena · 2026-10-03
- IFM open-sources K2-Type-0.9B: a decision model outputting calibrated probabilities in one forward pass — HongyiWang10 · 2026-10-03
- Ex-Meta researcher joins Prime Intellect to build open-source frontier model INTELLECT-4 — willcb · 2026-10-03
- Custom profile pictures are rolling out to Claude apps — testingcatalog · 2026-10-03
- GPT-6.1 Sol hits SOTA on URSA retrosynthesis benchmark with 35% of molecules solved — DeryaTR_ · 2026-10-03
- Running 262K context on a 5090+4070 rig: three Strata patches hit 130 tok/s decode — Fz1zz · 2026-10-03