409k approval decisions: humans let ~1 in 3 malicious agent commands through

No-Conflict4823 · reddit · 2026-10-02

The author challenges the standard human-approval model for risky agent actions with data and proposed fixes:

The data

Proposed mechanisms

The author asks practitioners whether approval fatigue is real, which parts they'd enable, whether canaries would feel like surveillance, and what actually works today.

Original post →

More from coding & agent

coding & agent channel →