Evals are the new PRD, but they still can't tell you if you built the right thing
rseroter · x · 2026-09-09
Jeff Gothelf argues evals act as requirements docs — a definition of done proving the system works as expected — but not whether you built the right thing. He cites Albertsons data showing shoppers spend 10-26% more with its AI assistant (because they 'stop forgetting items'), a human outcome evals can't measure. As OpenAI CPO Kevin Weil and others push 'evals are the new PRD', the piece calls for a separate definition of done for AI features covering user behavior.
More from AGI Musings
- Alignment is ill-defined, but monitorability and reversibility aren't — Afinetheorem · 2026-09-09
- Founder's honest take: bullish on AI, anxious about the pace and opacity — adityaag · 2026-09-09
- a16z: consumer AI is at the 2010 iPhone moment — the giants haven't been built yet — lennysan · 2026-09-09
- Terence Tao posts insightful thread reacting to AI's Navier-Stokes claim — fchollet · 2026-09-09
- Scooped Researchers Publish Less in Top Journals and Get 21% Fewer Citations — soumitrashukla9 · 2026-09-09
- Newton with 150 Hires? OpenAI Scooping Row Draws Leibniz Analogy — JMannhart · 2026-09-09