409k approval decisions: humans let ~1 in 3 malicious agent commands through
No-Conflict4823 · reddit · 2026-10-02
The author challenges the standard human-approval model for risky agent actions with data and proposed fixes:
The data
- A browser game by Alex Wauters at ScaleX puts players in the approver seat for an AI coding agent: across 40k+ plays and 409k decisions, about 1 in 3 malicious commands slip through. The most-missed one was approved 65% of the time — even though the log right above showed what it did.
- Analogy: published studies show doctors override anywhere from half to nearly all drug safety alerts. Stakes make people read; volume makes them stop.
Proposed mechanisms
- Blind second review: an AI reviewer independently judges each risky action; the human doesn't see its opinion first.
- Disagreement escalation: human approves something the AI flagged → a second person signs off.
- Canaries: TSA-style — synthetic "obviously deny" actions injected into the queue (logged, never executed) to check who's actually reading.
- Earned approval authority: reviewers build accuracy records; too many missed canaries narrows your approval scope.
The author asks practitioners whether approval fatigue is real, which parts they'd enable, whether canaries would feel like surveillance, and what actually works today.
More from coding & agent
- CloudAnalyzer: browser-based LiDAR tools for SLAM map fixes and Lanelet2 editing in Rust+WASM — rsasaki0109 · 2026-10-02
- AI Coding Accelerates PS5 Emulation: SharpEmu Now Runs 10 Playable Titles — mark_k · 2026-10-02
- Muse adds a main agent role to help users figure out what to automate — op7418 · 2026-10-02
- shardr: a content-addressed model store with BitTorrent sync and OpenAI-compatible serving — Cyb3erDudu · 2026-10-02
- Cloudflare adds Analytics SQL binding to query analytics datasets from Workers — ritakozlov · 2026-10-02
- a16z podcast: Lio CEO on why AI-native startups win work outside the system of record — a16z · 2026-10-02