Turning coding-agent failures into regression tests with Kitaru session replay
strickvl · x · 2026-09-29
- The author wants every coding-agent failure to make the system harder to break, so they integrated Kitaru (by ZenML) into their Decode project.
- The workflow: every run is recorded as an inspectable session, similar failures are grouped, and failures become regression tests.
- When changing model, prompt, tools, or harness, those exact sessions can be replayed against the new version to check regressions—a reusable recipe for agent eval engineering.
More from coding & agent
- We killed lines-of-code metrics, so why are we measuring AI productivity by merged PRs? — brandon_galang · 2026-09-29
- Every's new skill "Is This Anything" turns your unfinished AI experiments into lessons — every · 2026-09-29
- LlamaIndex explores on-the-fly model routing for document parsing tasks — llama_index · 2026-09-29
- Steal this prompt: Opus 5.5 plus Runway MCP for polished motion design videos — notiansans · 2026-09-29
- Pluto launches in beta: a personal agent with inbox, memory and a computer — Rasmic · 2026-09-29
- DeepLedger Brings MCP to QuickBooks with Review Tasks and Company Memory — Sorry-Clothes931 · 2026-09-29