Do We Use Bootstrap for Agent Evaluations Too?
abcdedcbaa · reddit · 2026-07-19
The post asks: When conducting agent evaluations, should we use large test sets plus statistical sampling, similar to chatbot evals?
The author's current practice is to evaluate with about 70 golden cases, then expand to around 1k+ cases via bootstrap / resampling as a sanity check. They want to know if others do the same for agent scenarios, and whether current agent eval / observability standards have shifted towards trajectory-level metrics, such as tool call sequences, reasoning processes, and final answers.
More from coding & agent
- Codex helps build Valdiluce, an open-world game with climbing, gliding and gondolas — Dimillian · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- LangSmith adds tracing for Pipecat, LiveKit, OpenAI Realtime, and Gemini Live — LangChain · 2026-07-22
- An MCP server signs every AI agent tool call into a verifiable Merkle chain — Funky_Chicken_22 · 2026-07-22
- Annotated transcript of a Claude Code team interview is now available — trq212 · 2026-07-22