Human Oversight Fails: Only 5% Catch Dangerous AI Commands After 50 Prompts

dhadfieldmenell · x · 2026-08-12

Anthropic conducted an experiment with 1,053 paid testers to evaluate the effectiveness of a "human in the loop." They secretly swapped a permission prompt with a clearly dangerous text command.

Testers caught the dangerous command only 13.6% of the time. After 50 prompts, human detection dropped to around 5% due to alert fatigue. In contrast, the AI's auto-mode successfully blocked 89% of the same commands, remaining flat across session length.

The findings challenge the concept of mandatory human oversight: if humans blindly approve prompts, supervision becomes ineffective, making automated AI safety mechanisms significantly more reliable.

Original post →

More from coding & agent

coding & agent channel →