Replaying agent sessions with cached tool outputs to test fixes and model swaps safely
pauliusztin · reddit · 2026-10-03
The hardest part of agent regression testing is recreating the exact failure state. The author's answer: record every tool call, then replay sessions.
- Replay ≠ playback: the agent runs from scratch with your new code/prompt/model, but tool calls matching the recording get cached outputs; new arguments or changed tools run for real.
- Four per-tool replay policies: history (cache), static (mock), passthrough (real), llm (model invents result); uncached tools must passthrough or fail. Crashed runs need passthrough and are only safe inside a container.
- Workflow: replay unchanged to prove the recording reproduces the bug, fix, replay again, diff. Model swaps work the same — he overrode every Sonnet call with gemini-3.8-flash and replayed a 4-session failure group, all green.
- Caveats: replays burn real tokens, keep groups small; drop groups that haven't failed recently. He used Kitaru as the record/replay layer on a Pydantic AI agent.
More from coding & agent
- Lucy hits 94.7% Recall@10 on FinanceBench with Qdrant over 2.8M SEC/DART filings — qdrant_engine · 2026-10-03
- The AI slop loop: agents rewriting each other's junk, and thin context is the root cause — IgorCarron · 2026-10-03
- 8 core networking concepts every developer should understand — goyalshaliniuk · 2026-10-03
- Codex vs Claude Opus head-to-head: porting retro console games from CPU to GPU rendering — ssh4net · 2026-10-03
- Open-source iFixAi audits AI agents in 120s with 60 checks and an A-F grade — thisdudelikesAI · 2026-10-03
- Git doesn't fit parallel agents: worktrees eat disk, so agents will design the alternative — Al_Grigor · 2026-10-03