cua-speedrun: the time-performance Pareto frontier is split across model families
kohjingyu · x · 2026-10-01
The cua-speedrun team notes that models and reasoning-effort settings carry different performance/speed/cost tradeoffs: on OSWorld-Verified, the time-performance Pareto frontier is dominated by a mix of families and settings, including Opus 5.5, GPT-6 Astra, and GPT-5.6 Luna.
Related event: cua-speedrun Goes Open Source: Leaderboard, Paper, and Code Released(2 posts)→
More from Research
- AIDE² paper shows AI research agent recursively rewriting its own code, 7 gains in 8-day run — krishnan · 2026-10-02
- CRUG lets a single RNN learn new dynamical systems with zero forgetting by recycling units — tweetsatpreet · 2026-10-02
- RL inside the harnesses: lifting LFM2.5 from 42% to 54% across four agent harnesses — _lewtun · 2026-10-02
- Kyutai's 100M-parameter on-device PocketTTS is first speech model trained with drifting — kastnerkyle · 2026-10-02
- PostTrainBench v1.2: Fable 5.1 takes #1 at 44.6%, Opus 5.5 second, now reproducible via Harbor — dejavucoder · 2026-10-02
- Injectable nanoparticle "solar cells" restore light sensitivity in blind retinas — DrKavner · 2026-10-02