Evaluating Agents: Real Output is State Change, Not Natural Language
Harshit-24 · reddit · 2026-07-31
The author argues that for AI agents with tool-calling capabilities, evaluating their effectiveness requires looking beyond natural language responses to focus on actual state changes in external systems.
Using the Komo AI team's practice as an example, a robust agent audit mechanism should not only log the final text but also include: the tool called, the exact operation and target, the system state before and after the operation, human approval records, and potential failures. The author recommends staging external communications for human review and keeping the system of record outside the chat context, preventing chat history from becoming the sole proof of action.
More from coding & agent
- Three Claude Skills Every Engineering Org Should Build — arpit_bhayani · 2026-07-31
- 10k-Star Reverse-Skill: AI Routing Pack for Pentesting in Coding Agents — zhaoxuya520 · 2026-07-31
- Benchmarking Kimi K3, GLM 5.2, and DeepSeek V4 Pro in Agent Workflows — Teknium · 2026-07-31
- Open-Source KG Extractor Runs Qwen on a Single NVIDIA L4 — JeremyCMorgan · 2026-07-31
- Anthropic Engineers Run Hundreds of Agents via Graph Engineering — blaizedsouza · 2026-07-31
- Build and Deploy Websites from Scratch Using Claude Code and MCP Connectors — TawohAwa · 2026-07-31