Human Oversight Fails: Only 5% Catch Dangerous AI Commands After 50 Prompts
dhadfieldmenell · x · 2026-08-12
Anthropic conducted an experiment with 1,053 paid testers to evaluate the effectiveness of a "human in the loop." They secretly swapped a permission prompt with a clearly dangerous text command.
Testers caught the dangerous command only 13.6% of the time. After 50 prompts, human detection dropped to around 5% due to alert fatigue. In contrast, the AI's auto-mode successfully blocked 89% of the same commands, remaining flat across session length.
The findings challenge the concept of mandatory human oversight: if humans blindly approve prompts, supervision becomes ineffective, making automated AI safety mechanisms significantly more reliable.
More from coding & agent
- AI Kills the 'Bus Factor', Reshaping Software Evaluation — Elijah_Meeks · 2026-08-12
- oh-my-pi v17.2.14 Released: Introduces External Thinking and Disable Reasoning Options — banteg · 2026-08-12
- Sourcegraph: AI-Generated Code Exacerbates Cross-Repo Security Vulnerabilities — glenbeer · 2026-08-12
- Microsoft Research: Injecting NL Skills Cuts Agent Reasoning Cost by 6x — dair_ai · 2026-08-12
- Codex CLI Update: `resume` and `fork` Now Load Instantly — charliermarsh · 2026-08-12
- Going All-In on AI Coding: A Year of Reflection on the Death of Manual Implementation — UAP44 · 2026-08-12