Hamel Husain: make AI outputs easier to evaluate by redesigning the product, not the evals
HamelHusain · x · 2026-09-26
Hamel Husain and Shreya Shankar answer "how can I make AI outputs easier for people to evaluate?" — start with product design, not the eval pipeline.
- Surface intermediate outputs: instead of asking a doctor to critique a generated medical report, show the extracted facts with links to sources and let them correct facts or resolve conflicting evidence before generating the report — keeping the human in the loop and building trust
- Reduce review friction: render outputs in familiar formats (emails as emails, syntax-highlighted code), keep reviewer context on one screen with collapsible details, add keyboard shortcuts for moving between examples and recording judgments, and show progress like "45 of 100 examples reviewed"
The companion post "It's Hard to Eval" Is a Product Smell expands on this with before-and-after mockups.
More from coding & agent
- Runway MCP lands in ChatGPT plugin directory, generating full ads from a single prompt — runwayml · 2026-09-26
- Runway MCP brings Gen-4.5 video and image generation into Claude chats — runwayml · 2026-09-26
- Anthropic engineer: Claude Tag writes >50% of my PRs every day — bcherny · 2026-09-26
- Cline Desktop adds SSH support: agent runs on remote machines, UI stays local — cpaik · 2026-09-26
- Nautilo: open-source multi-user platform where humans and AI agents share one terminal session — Dan_Jeffries1 · 2026-09-26
- Dev builds 8 Astra skills automating the full creator-marketing pipeline, end to end — alexgoughcooper · 2026-09-26