Evals are the new PRD, but they still can't tell you if you built the right thing

rseroter · x · 2026-09-09

Jeff Gothelf argues evals act as requirements docs — a definition of done proving the system works as expected — but not whether you built the right thing. He cites Albertsons data showing shoppers spend 10-26% more with its AI assistant (because they 'stop forgetting items'), a human outcome evals can't measure. As OpenAI CPO Kevin Weil and others push 'evals are the new PRD', the piece calls for a separate definition of done for AI features covering user behavior.

Original post →

More from AGI Musings

AGI Musings channel →