AI Evals FAQ grows to 48 Q&As: sensitive data, huge traces, and stale gold datasets

HamelHusain · x · 2026-09-16

Hamel Husain and Shreya Shankar updated their AI Evals FAQ with 5 new questions, bringing it to 48 Q&As. The guide distills questions from their course that has trained 700+ engineers and PMs, with deliberately sharp, opinionated answers.

New entries cover: whether you need a reference answer or rubric before annotating data, doing evals when traces contain sensitive data, reviewing very large traces, how much context to give an LLM judge, and what to do when your gold eval dataset goes stale.

The doc also cleanly separates model benchmarks (GPQA Diamond, Terminal-Bench, MMLU) from product evals, stressing that most AI products need multiple evals because they fail in different ways. A high-density reference for anyone doing LLM eval engineering.

Original post →

More from coding & agent

coding & agent channel →