Advanced AI evals: most teams skip error discovery and measure the wrong things

lennysan · x · 2026-09-23

In a deep-dive Lenny's Newsletter post, Hamel Husain and Shreya Shankar distill work with 50+ AI companies into a guide for finding and fixing hidden AI failures. Their core claim: most teams jump straight to writing metrics and end up measuring the wrong things — the skipped step is error discovery.

Case data: Ramp lifted automatic receipt-collection accuracy from 35% to 83%; Shopify shipped an AI workflow builder 2.2x faster and 68% cheaper than the frontier-model setup it replaced; Harvey nearly doubled internal quality on its rebuilt AI contract reviewer; Cursor tuned Auto Balance routing for higher satisfaction at 41% lower cost.

The post details which steps can and cannot be automated, plus a free plugin that lets a coding agent do most of the heavy lifting, alongside a 30-minute read and ready-to-use evals skills. Nearly half of recent AI PM job openings ask for evals experience.

Original post →

More from coding & agent

coding & agent channel →