Do We Use Bootstrap for Agent Evaluations Too?
abcdedcbaa · reddit · 2026-07-19
The post asks: When conducting agent evaluations, should we use large test sets plus statistical sampling, similar to chatbot evals?
The author's current practice is to evaluate with about 70 golden cases, then expand to around 1k+ cases via bootstrap / resampling as a sanity check. They want to know if others do the same for agent scenarios, and whether current agent eval / observability standards have shifted towards trajectory-level metrics, such as tool call sequences, reasoning processes, and final answers.
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11