ApprenticeBench: Top Models Now Beat APIs Through GUIs, the 'CUA Tax' Has Disappeared

ysu_nlp · x · 2026-09-11

Boyu Gou (SeeAct, UGround, Mind2Web 2) launched ApprenticeBench, the first benchmark for computer use and continual learning in a realistic accounting job. On a 100-task run combining learning and execution, Fable 5.1 scored 72% via GUI vs 70% via API, and GPT-6 Astra 68% vs 65%—meaning the 'CUA tax' has disappeared for top models, with GUI-based agents now outperforming API workflows. He argues computer use has been seriously underestimated and evaluated too narrowly since Opus 4 and Sonnet 4.5.

Related event: NeoCognition Releases ApprenticeBench, First Job-Level Continual Learning Benchmark(8 posts)→

Original post →

More from coding & agent

coding & agent channel →