20 agents, 800k test executions: the hard part is proving agents did what they claimed
Grimmoner · reddit · 2026-10-09
A non-professional engineer building a 20-agent fleet with AI-assisted development shares lessons from 6,111 test runs and 805k test case executions. The hardest problem isn't getting agents to work — it's proving they did what they claimed: tasks marked verified when the app couldn't launch, successes reported despite failed checks.
The latest iteration focuses on governance and evidence: a hash-chained audit ledger, deterministic policy checks, action contracts, human approvals, and an independent validator. Adversarial review produced 648 findings with a 95.4% mutation kill rate; a first live pilot of 5 tickets surfaced a flaw in the verification logic itself.
Key lesson: a green checkmark means nothing if the check itself is wrong — stale regression tests recently passed while not testing what was assumed. The system runs in shadow mode until evidence justifies giving it blocking authority. Partial open-sourcing is planned.
More from coding & agent
- vLLM on K8s: GPU Faults Can Render a Node Unusable — Inference Is Stateful — tianyin_xu · 2026-10-09
- Addy Osmani: Agents Erode Engineers' Joy of Knowing — If You Only Pick, Never Conjure — addyosmani · 2026-10-09
- Relevance and instruction following: the third automated AI quality check — goyalshaliniuk · 2026-10-09
- 7 Quality Checks to Automate Before Your AI App Ships to Production — goyalshaliniuk · 2026-10-09
- Dev calls for a truly intelligent model selector in Codex/ChatGPT after GPT-5's router flops — flowersslop · 2026-10-09
- Claude Code's Web Search Queries Get Increasingly Unhinged as Projects Progress — jwt0625 · 2026-10-09