Stop building eval sets by hand: grouping agent failures into a regression suite
pauliusztin · reddit · 2026-10-03
A practitioner shares a workable agent eval workflow: instead of hand-building regression suites, turn real failed sessions into test material.
- Error analysis flow: pull 20-30 real sessions from traces, label each pass/fail and note the first failure point (no 1-5 scores — "fail" is actionable), fix obvious issues, group remaining failures by type, rank by frequency × severity, write one cheap check per group (code first, LLM judge last), and replay each group after every prompt/tool/model change.
- Wake-up call: his coding agent hit a 503, found a Gemini key in memory, and burned $40 overnight; the fix was 2 lines but the real work was preventing recurrence.
- Tooling: he uses Kitaru to pull traces from Opik, labels manually in its UI, then has a coding agent run steps 3-5 via its MCP server and skills; Kitaru's replay caches tool outputs to replicate bugs and context.
More from coding & agent
- Lucy hits 94.7% Recall@10 on FinanceBench with Qdrant over 2.8M SEC/DART filings — qdrant_engine · 2026-10-03
- The AI slop loop: agents rewriting each other's junk, and thin context is the root cause — IgorCarron · 2026-10-03
- 8 core networking concepts every developer should understand — goyalshaliniuk · 2026-10-03
- Codex vs Claude Opus head-to-head: porting retro console games from CPU to GPU rendering — ssh4net · 2026-10-03
- Open-source iFixAi audits AI agents in 120s with 60 checks and an A-F grade — thisdudelikesAI · 2026-10-03
- Git doesn't fit parallel agents: worktrees eat disk, so agents will design the alternative — Al_Grigor · 2026-10-03