How Comet built Opik Diagnostics: automating agent debugging beyond the single trace
techNmak · x · 2026-10-03
Comet's engineering team published a detailed retrospective on building Diagnostics, their automated agent debugging tool for the Opik observability platform — including two false starts. The problem: manual workflows (find a trace, paste the JSON into Claude Code) don't scale to thousands of daily traces, and the most damaging failures never throw exceptions — agents quietly retry failing tools or drift in behavior no one notices. The article covers how they designed automated diagnostics, several approaches to having models inspect large numbers of traces, and the final decision to move much of large-scale checking into ordinary queries over trace data.
More from coding & agent
- AI Village dataset with millions of agent behavior samples trends on Hugging Face — aidigestorg · 2026-10-03
- NYC Agentic AI meetup to recap September AI moves with live PowerShell decision demo — dfinke · 2026-10-03
- sindresorhus: AI-made PRs make humans mere routers — open source should let project AIs absorb contributions directly — vykthur · 2026-10-03
- Willowmere v2: Claude-coded 3D pixel art game ships with zero asset files, pure HTML/JS — BroEvenIDK · 2026-10-03
- Five prompts that took Every's ops lead from one-off chats to delegating projects to agent teams — danshipper · 2026-10-03
- AI engineering is like making law, not playing games: rules must shift as they meet reality — danshipper · 2026-10-03