How Do You Regression-Test Agents When Every Run Takes a Different Path?
No-Age-3362 · reddit · 2026-09-30
A Reddit discussion on regression-testing agents: unit tests assume deterministic input-output, but agents call three tools on one run and five on the next, both correct. Worse, one run can silently skip a required check while still looking fine.
- Comparing final outputs misses silent skipped checks
- Comparing full trajectories flags legitimate multi-path solutions as failures
Partial approaches covered: outcome-only assertions, asserting a few required/forbidden steps (must call X before Y, never call Z), scoring pass rates over multiple runs, and LLM-as-judge on traces. The author asks what teams actually use in production and how many runs per case before trusting a pass rate.
More from coding & agent
- Constella agent drives Blender via MCP: rough edges, high upside — sidahuj · 2026-09-30
- Law firm AI working group tests multiple models — Claude keeps outperforming — jkubicki · 2026-09-30
- Hands-on with Dot: smarter than Grok Bot but keeps ignoring custom skill constraints — jdjohnson · 2026-09-30
- OpenAI launches cloud API for testing web pages; independent QA seen as startup moat — hugs · 2026-09-30
- Claude+small-model harness cuts 1,000 AI decisions from $605 to $0.17 — jamestagg · 2026-09-30
- How do you evaluate the quality of AI agent-generated long-form writing? — OwlZealousideal4779 · 2026-09-30