Hamel Husain: "It's hard to eval" is a product smell, not just an eval problem

hugobowne · x · 2026-09-01

Hugo Bowne-Anderson will host a livestream with AI evals expert Hamel Husain on "Stop Shipping AI Nobody Can Verify."

Core thesis: when an AI data agent reports net revenue of $4.21M without showing the metric definition, query, assumptions, or what it couldn't verify, users must redo the analysis to trust it—that's a product problem. Hamel argues "hard to eval" is a product smell: design for verification first, letting users inspect evidence, compare against trusted references, and review smaller units of work.

Topics include studying the checks domain experts perform, why answers without provenance force users to redo the agent's work, what data agents should expose (metric definitions, intermediate calculations, source queries, sanity checks), progressive disclosure, breaking output into accept/edit/reject units, and how verification-friendly design yields cheaper annotation and stronger eval signals.

Related event: Hamel Husain Talks AI Verification: Hard-to-Evaluate Agents Are a Product Flaw(2 posts)→

Original post →

More from coding & agent

coding & agent channel →