Demo: Automating AI Evaluations the Right Way
HamelHusain · x · 2026-07-07
Hamel Husain and Shreya release a new episode on how to automate AI evaluations correctly. The core idea is that finding issues is the most critical part of the eval workflow, and Shreya demonstrates how to iteratively prompt AI to uncover "unknown unknowns." Topics covered include why vendors want to automate evals for you, why no tool can fully automate the process, common mistakes (like directly asking AI to "find problems"), proper AI-assisted error analysis, building a review interface from scratch, and labeling traces.
More from coding & agent
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11