Fail an AI agent twice, add regression tests: a practical agent debugging workflow
strickvl · x · 2026-10-08
- Developer Paul Iusztin shares the rule he follows while building Decode, a coding agent: fail once, blame the agent; fail twice, add a regression test. The point isn't that agents fail, but what you do with each failure.
His debugging workflow:
- Record agent runs as sessions; cluster similar failures and prioritize by frequency × severity
- Trace important failures through the session to find the root cause
- Write one evaluator per failure class answering a single question: did this failure happen again?
- Fix → replay → compare: change a prompt, tool, model, skill, or harness piece, then replay the same session to verify the fix
He uses Kitaru by ZenML, which records runs as sessions, groups recurring failures into cohorts, and can reuse recorded tool outputs when identical tool calls recur—only changed calls re-execute—so you don't have to painfully recreate the exact context that caused a failure.
More from coding & agent
- Agents love tidying up files nobody asked them to touch — JFPuget · 2026-10-08
- Open-source bridge turns a browser chat tab into an OpenAI-compatible API endpoint — harshanacz · 2026-10-08
- GraphRAG Bug: Deleted Documents Stay Indexed and Retrievable After Updates — JeremyCMorgan · 2026-10-08
- Paper: Vibe Coding Kills Open Source as AI-Recommended Repos Lose Stars — soumitrashukla9 · 2026-10-08
- Every publishes definitive guide to Compound Engineering, the philosophy behind 7k-star plugin — every · 2026-10-08
- Claude Haiku 5.5 matches GPT-6 Luna pricing but a stingier tokenizer hides a 1.25x cost hike — Simon Willison · 2026-10-08