Evaluating Agents: Real Output is State Change, Not Natural Language

Harshit-24 · reddit · 2026-07-31

The author argues that for AI agents with tool-calling capabilities, evaluating their effectiveness requires looking beyond natural language responses to focus on actual state changes in external systems.

Using the Komo AI team's practice as an example, a robust agent audit mechanism should not only log the final text but also include: the tool called, the exact operation and target, the system state before and after the operation, human approval records, and potential failures. The author recommends staging external communications for human review and keeping the system of record outside the chat context, preventing chat history from becoming the sole proof of action.

Original post →

More from coding & agent

coding & agent channel →