Humans Missed 1 in 3 Threats When Approving AI Agent Commands Across 40,000 Plays
Wirbelwind · reddit · 2026-08-06
A browser game testing 'human-in-the-loop' security for AI coding agents collected over 400,000 approve/deny decisions. Analysis reveals that human players missed about 1/3 of malicious threats.
Key Data Insights:
- Credential Theft: cat /.ssh/idrsa was blocked by 82% of players, but other sensitive config and credential files were waved through about half the time.
- Disguised Execution: If malicious code modified package.json and requested to run as an npm run command (e.g., npm run analyze), it was approved 65% of the time, even with the evil payload explicitly visible in the execution history log.
This indicates that current human review mechanisms are highly fragile against carefully disguised agent commands.
More from coding & agent
- Rumor: Alibaba Ships 2.4T Parameter Model for Autonomous CLI Dev — eyishazyer · 2026-08-06
- Vocci Demo: Turns Spoken Discussions into Claude Prototypes and Code — Div_pradeep · 2026-08-06
- Qwen3.8-Max Built a CLI Tool Autonomously for 16 Days Straight — socialwithaayan · 2026-08-06
- Qwen3.8-Max Tested: Builds CLI Tool Autonomously for 16 Days, Huge Leap in Long-Horizon Tasks — socialwithaayan · 2026-08-06
- MiniMax Video Workflow Tip: Tweaking Reference Node Settings for Better Face Likeness — xDFINx · 2026-08-06
- Voice + Agents Guide: Turning Spoken Half-Thoughts Directly into Code — every · 2026-08-06