Do We Use Bootstrap for Agent Evaluations Too?

abcdedcbaa · reddit · 2026-07-19

The post asks: When conducting agent evaluations, should we use large test sets plus statistical sampling, similar to chatbot evals?

The author's current practice is to evaluate with about 70 golden cases, then expand to around 1k+ cases via bootstrap / resampling as a sanity check. They want to know if others do the same for agent scenarios, and whether current agent eval / observability standards have shifted towards trajectory-level metrics, such as tool call sequences, reasoning processes, and final answers.

Original post →

More from coding & agent

coding & agent channel →