cua-speedrun: the time-performance Pareto frontier is split across model families

kohjingyu · x · 2026-10-01

The cua-speedrun team notes that models and reasoning-effort settings carry different performance/speed/cost tradeoffs: on OSWorld-Verified, the time-performance Pareto frontier is dominated by a mix of families and settings, including Opus 5.5, GPT-6 Astra, and GPT-5.6 Luna.

Related event: cua-speedrun Goes Open Source: Leaderboard, Paper, and Code Released(2 posts)→

Original post →

More from Research

Research channel →