ZenML open-sources Kitaru: replay production agent traces against different models or prompts

htahir1 · reddit · 2026-08-18

ZenML open-sourced Kitaru (Apache 2.0, fully self-hostable) to answer: given a real production agent session, what would have happened if you changed the model, prompt, context, or tool policy — without calling live tools again.

Key mechanics:

The broader workflow: production traces → investigate → cohorts → expert judgment → evaluators → replay → experiments. Kitaru proposes cohorts of similar sessions from recurring patterns; expert reviews calibrate evaluators, which are Python functions versioned alongside the agent.

The maintainer seeks feedback on two questions: how to handle tool calls during replay, and what replay fidelity is required before trusting a counterfactual.

Related event: ZenML Open-Sources Kitaru for Replay-Based Agent Evaluation(2 posts)→

Original post →

More from coding & agent

coding & agent channel →