Dev builds agent testing tool that catches fake tool calls and loops in full conversations

mbtigeekjung · reddit · 2026-10-01

A developer shared a tool that runs agents through full conversations and flags failure patterns: re-asking questions already answered, claiming an action happened without calling the tool, and losing customers ready to convert.

The post also asks the community three practical questions: test full conversations or single responses? How to catch agents saying they did something they didn't? Real transcripts, simulated users, or vibes? Useful context for anyone working on agent eval engineering.

Original post →

More from coding & agent

coding & agent channel →