Do We Use Bootstrap for Agent Evaluations Too?
abcdedcbaa · reddit · 2026-07-19
The post asks: When conducting agent evaluations, should we use large test sets plus statistical sampling, similar to chatbot evals?
The author's current practice is to evaluate with about 70 golden cases, then expand to around 1k+ cases via bootstrap / resampling as a sanity check. They want to know if others do the same for agent scenarios, and whether current agent eval / observability standards have shifted towards trajectory-level metrics, such as tool call sequences, reasoning processes, and final answers.
More from coding & agent
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11