a16z Report: Computer-Using Agents Now Beat Human Testers

a16z · x · 2026-08-10

a16z published an in-depth analysis on Computer-Using Agents (CUAs). Data shows that CUA models have improved faster than expected over the past year: the best model scored 85% on the standard desktop benchmark, up from 42% last year, surpassing the human tester average of 72%.

The article notes that CUAs have crossed from demo to deployable in production. They are now holding up at scale on narrow, repeatable workflows like updating systems of record, moving data through portals, and processing tickets. However, agents remain brittle when work drifts off the runbook.

Related event: a16z Report: Computer-Use Agents Surpass Humans in Benchmarks(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →