Agents write 741% more code but ship only 30% more software — human review is the new bottleneck

AI Engineer · youtube · 2026-09-30

Laurie Voss (Arize AI, npm co-founder) reviews the data on code review in the agent era: developers using autonomous agents wrote 741% more code but shipped only 30% more software, and reviewer effectiveness collapses past 400 lines while agents open 10,000-line PRs. Evidence includes OpenAI's zero-human-code product, METR's finding that about half of SWE-bench-passing PRs wouldn't be merged, and Cognition's FrontierCode scoring 88% on SWE-bench Pro vs. 29% real mergeability — a mergeability benchmark would instantly become a training signal. She covers how Cursor and GitHub run production review, why multi-pass review and default suspicion cut false positives, and what Carlini's agent-built C compiler and Bun's million-line Zig-to-Rust port (13,044 unsafe blocks) say about removing humans. Automated reviewers can be fooled by prompt injection, leaving production as the last reviewer standing. Code review is being rebuilt, not killed — build a review harness today.

Related event: Code Review Becomes the Bottleneck as Agent Coding Surges 741%(2 posts)→

Original post →

More from coding & agent

coding & agent channel →