Auditing 1,228 Human Interventions: 91% Weren't Decisions, So They Stopped Trusting Agent "Done"

kraboo_team · reddit · 2026-09-28

A team running 25 AI agent workspaces (PM → Dev → Tester → Critic → Release) audited 8 weeks of human interventions and found only 9% were real decisions; 59% of tasks were closed by hand despite agents claiming completion, and 8-11 per 100 tasks were false "done" claims (e.g., "30 files pushed" with nothing on the remote).

Their fix: keep agents as-is but add a small MCP-backed verification service:

They're starting in shadow mode and asking the community about measuring verifier miss rates, subjective checks without LLM judges, and continuous re-verification after "done".

Original post →

More from coding & agent

coding & agent channel →