Hamel Husain: "It's hard to eval" is a product smell, not just an eval problem
hugobowne · x · 2026-09-01
Hugo Bowne-Anderson will host a livestream with AI evals expert Hamel Husain on "Stop Shipping AI Nobody Can Verify."
Core thesis: when an AI data agent reports net revenue of $4.21M without showing the metric definition, query, assumptions, or what it couldn't verify, users must redo the analysis to trust it—that's a product problem. Hamel argues "hard to eval" is a product smell: design for verification first, letting users inspect evidence, compare against trusted references, and review smaller units of work.
Topics include studying the checks domain experts perform, why answers without provenance force users to redo the agent's work, what data agents should expose (metric definitions, intermediate calculations, source queries, sanity checks), progressive disclosure, breaking output into accept/edit/reject units, and how verification-friendly design yields cheaper annotation and stronger eval signals.
More from coding & agent
- 实测:Qwen 27B 在配置调试中一次性修复了 DSV4 无法解决的问题 — OvertaxedOne · 2026-09-01
- AI Agents现状:效果惊艳但仍需大量人工干预 — Hot-Cup3451 · 2026-09-01
- SOTA agents generate useless code; we need precise definitions for bad code — alexisgallagher · 2026-09-01
- Dev shares HTML to PPTX solution using pptxgenjs — dotey · 2026-09-01
- Developer Observes Persistent Dumb Decisions in AI Code Outputs — rickasaurus · 2026-09-01
- Inside Anthropic's Hackathon: Can AI One-Shot a Full Unity 3D Game? — davidfromkansas · 2026-09-01