AI Evals FAQ grows to 48 Q&As: sensitive data, huge traces, and stale gold datasets
HamelHusain · x · 2026-09-16
Hamel Husain and Shreya Shankar updated their AI Evals FAQ with 5 new questions, bringing it to 48 Q&As. The guide distills questions from their course that has trained 700+ engineers and PMs, with deliberately sharp, opinionated answers.
New entries cover: whether you need a reference answer or rubric before annotating data, doing evals when traces contain sensitive data, reviewing very large traces, how much context to give an LLM judge, and what to do when your gold eval dataset goes stale.
The doc also cleanly separates model benchmarks (GPQA Diamond, Terminal-Bench, MMLU) from product evals, stressing that most AI products need multiple evals because they fail in different ways. A high-density reference for anyone doing LLM eval engineering.
More from coding & agent
- Same Prompt, Four Models Behind One MCP: Only One Got It Right — rohanpaul_ai · 2026-09-16
- Shumer: Skip Complex Setups, One Manager Agent Session Is Enough — mattshumer_ · 2026-09-16
- Google Cloud API Gateway now acts as a remote MCP server for existing REST APIs — rseroter · 2026-09-16
- Dev ports Blender 5.1 MCP add-on to Blender 3.6, AI builds a full owl scene from one prompt — SpinachOk9137 · 2026-09-16
- Lyft cut support agent ship time from 6 months to 1-2 weeks with LangGraph and LangSmith — LangChain · 2026-09-16
- Text-only AI agent beats Doom at ~10 calls/sec, costing about $7 per hour — hackgoofer · 2026-09-16