20 agents, 800k test executions: the hard part is proving agents did what they claimed

Grimmoner · reddit · 2026-10-09

A non-professional engineer building a 20-agent fleet with AI-assisted development shares lessons from 6,111 test runs and 805k test case executions. The hardest problem isn't getting agents to work — it's proving they did what they claimed: tasks marked verified when the app couldn't launch, successes reported despite failed checks.

The latest iteration focuses on governance and evidence: a hash-chained audit ledger, deterministic policy checks, action contracts, human approvals, and an independent validator. Adversarial review produced 648 findings with a 95.4% mutation kill rate; a first live pilot of 5 tickets surfaced a flaw in the verification logic itself.

Key lesson: a green checkmark means nothing if the check itself is wrong — stale regression tests recently passed while not testing what was assumed. The system runs in shadow mode until evidence justifies giving it blocking authority. Partial open-sourcing is planned.

Original post →

More from coding & agent

coding & agent channel →