Advanced AI evals: most teams skip error discovery and measure the wrong things
lennysan · x · 2026-09-23
In a deep-dive Lenny's Newsletter post, Hamel Husain and Shreya Shankar distill work with 50+ AI companies into a guide for finding and fixing hidden AI failures. Their core claim: most teams jump straight to writing metrics and end up measuring the wrong things — the skipped step is error discovery.
Case data: Ramp lifted automatic receipt-collection accuracy from 35% to 83%; Shopify shipped an AI workflow builder 2.2x faster and 68% cheaper than the frontier-model setup it replaced; Harvey nearly doubled internal quality on its rebuilt AI contract reviewer; Cursor tuned Auto Balance routing for higher satisfaction at 41% lower cost.
The post details which steps can and cannot be automated, plus a free plugin that lets a coding agent do most of the heavy lifting, alongside a 30-minute read and ready-to-use evals skills. Nearly half of recent AI PM job openings ask for evals experience.
More from coding & agent
- LangSmith adds dedicated tracing UI for Jev-style decision model calls in agents — hwchase17 · 2026-09-23
- After a week with Opus 5.5: 6 practical tips for getting the most out of it — every · 2026-09-23
- Ant's Ling open-sources Ming-Image-0.1-Design, tops open-weight UI design leaderboard — bdsqlsz · 2026-09-23
- Horde: a local Rust daemon for durable, multi-agent coding orchestration via MCP — EyalToledano · 2026-09-23
- 12 Verified Listings on an Agent Hiring Marketplace, Zero Real Hires — Agent-OmegaLT · 2026-09-23
- After Burning Codex Limits in 3 Days, Developer Does Full Stack of Work with DeepSeek for $2 — MustafaAdam · 2026-09-23