How do you evaluate whether an agent change really improved anything?
Substantial_Step_351 · reddit · 2026-07-29
How do you tell whether an agent change actually improved anything when the same input can lead to different tool calls?
The poster asks whether people use a real eval setup for nondeterministic agent steps, or whether they mostly rerun the workflow a few times and inspect the outputs by hand.
More from coding & agent
- One GPU, many VMs: a practical guide to passthrough, vGPU, and API remoting — LoganGrasby · 2026-07-29
- Maven’s trending workshop teaches how to build AI agents from first principles — hugobowne · 2026-07-29
- Unreal Engine remains hard for agentic workflows because files, code and docs fight back — nptacek · 2026-07-29
- AI agents won’t replace you, but they can automate repetitive computer work — dotey · 2026-07-29
- LLM coding workflow adds a research gate to stop it from implementing every paper it finds — hypergraphr · 2026-07-29
- tldraw ships an offline app with Agent Skills for Claude and Codex drawing — vista8 · 2026-07-29