ApprenticeBench: Top Models Now Beat APIs Through GUIs, the 'CUA Tax' Has Disappeared
ysu_nlp · x · 2026-09-11
Boyu Gou (SeeAct, UGround, Mind2Web 2) launched ApprenticeBench, the first benchmark for computer use and continual learning in a realistic accounting job. On a 100-task run combining learning and execution, Fable 5.1 scored 72% via GUI vs 70% via API, and GPT-6 Astra 68% vs 65%—meaning the 'CUA tax' has disappeared for top models, with GUI-based agents now outperforming API workflows. He argues computer use has been seriously underestimated and evaluated too narrowly since Opus 4 and Sonnet 4.5.
More from coding & agent
- Ten lessons from three years building agents for real production work — garrytan · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Shopify CEO Tobi Lütke hails single-dev open-source agent harness Pi — aakashgupta · 2026-09-11
- Shopify CEO Tobi Lütke Calls Pi, a Solo-Dev Open-Source Agent Harness, 'the Most Interesting' One — aakashgupta · 2026-09-11
- Dev explains why MCP won him over: organic UX beats telling agents to run CLI commands — zeeg · 2026-09-11
- Sentry Founder: MCP Won Because Agents Can Just 'Fix This URL', Not Call CLIs — zeeg · 2026-09-11