ZenML launches Kitaru: Tinder-style swipe UI for evaluating agent traces
strickvl · x · 2026-09-03
ZenML launched Kitaru, a replay-based eval tool for AI agents that attacks a real UX gap: domain experts like lawyers or doctors will never log into an observability tool to annotate JSON dumps.
- Core idea: import production traces as replayable sessions, change one suspect thing, and re-run side by side to pinpoint failures.
- Tinder-style review: swipe through agent traces and leave voice notes on what went wrong, lowering the annotation barrier for non-technical reviewers.
- Workflow: wrap your agent in one line, import traces.l, cluster notes into cohorts, run experiments comparing versions (e.g. 90/90 failures down to 4/90), blind-validate your labels on held-out sessions, and compare model swaps (−61% cost with quality unchanged).
- Open on GitHub with 235 stars; 14-day free trial.
More from coding & agent
- Omnara: open-source, self-hostable alternative to Claude managed agents — JaynitMakwana · 2026-09-03
- NanoCodana: open-source Claude Code-style coding agent that runs entirely in the browser — andrepimentaa7 · 2026-09-03
- Solo user builds governance-first multi-agent system: lead agent, least privilege, independent auditor — Grimmoner · 2026-09-03
- Deploy the Foundry Model Router with Azure Bicep — adnan_hashmi · 2026-09-03
- Code Arena launches WebDev Pareto frontier to rank AI models by quality vs price — arena · 2026-09-03
- reverse-skill: 30k-Star GitHub repo routes AI agents through 44 reverse-engineering skill playbooks — alex_verem · 2026-09-03