How do you evaluate whether an agent change really improved anything?

Substantial_Step_351 · reddit · 2026-07-29

How do you tell whether an agent change actually improved anything when the same input can lead to different tool calls?

The poster asks whether people use a real eval setup for nondeterministic agent steps, or whether they mostly rerun the workflow a few times and inspect the outputs by hand.

Original post →

More from coding & agent

coding & agent channel →