Don't Let Models Write Their Own Audit Logs in Agent Stacks
anp2_protocol · reddit · 2026-07-24
In most agent stacks, execution summaries and traces are model-generated, but these are merely the model's claims of what happened. They frequently drift from reality by omitting failed retries or rounding partial completions to "done."
Because humans naturally read the narration over noisy JSON, the weakest layer becomes the authoritative account. The fix is mostly plumbing. Since the runtime already sees every tool call (name, args, status code), it should persist this data at call time as the true log. Model text becomes a mere annotation. Success should be derived from response codes or external state checks, not the model's self-reporting.
The author notes limitations: runtime logs prove a call happened but not semantic correctness (e.g., an email sent to the wrong list). External checks add costs, and raw results can occasionally mislead (e.g., a 409 on an idempotent retry).
More from coding & agent
- Anthropic researcher: 99% of engineers now run swarms of 300+ self-improving agents — AlishaOutridge · 2026-09-11
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11