A convincing agent trace can mislead: count occurrences, don't just read one

techNmak · x · 2026-10-03

A broken agent trace is dangerously convincing: repeated tool calls, bad arguments, thirty seconds spent on a five-second task. But the author argues that a trace only tells you what happened in one execution—not how common that behavior is across the system. It might happen twice in six months or one in five runs; the engineering priority is completely different.

The trendy fix—having another LLM read and summarize masses of traces (with a cheaper model if there are too many)—is often unnecessary. Once you know what you're looking for, the remaining questions are ordinary data questions: how often, which tool, since when, concentrated in which request type? These can be answered with plain queries over trace data—count failures, group them, pick representative examples, then bring in a model only when interpretation is needed. The lesson comes from Cometml's write-up on production agent diagnostics, where they moved large-scale checking from model inspection to ordinary queries over trace data.

Caveats: an agent can execute cleanly yet answer terribly; without a signal for that failure, counting won't reveal it—evaluation and careful trace reading still matter. Core takeaway: once you know what to look for, you don't need an LLM to count.

Original post →

More from coding & agent

coding & agent channel →