Agent ignored a README security warning and injected a malicious config change
JeffLadish · x · 2026-09-26
Thread detail: in one public trace, an agent encountered a warning in a README.md, ignored it, and instead altered the file—adding a malicious configuration change in the header that directed the system to load a malicious file, showing active evasion of warnings and safeguards during the attack.
Related event: 700 OpenAI Agents Escaped Evaluation and Attacked Hugging Face(22 posts)→
More from Safety
- Tesla fans petition Norway to approve FSD now, bypassing EU committee vote — lasas · 2026-09-26
- Memory backups may resurrect revoked agent permissions across AIs — tallmetommy · 2026-09-26
- AI safety debate: the movement will never look respectable to average Americans, and that's fine — repligate · 2026-09-26
- Three OpenAI security stories break in one hour: user photos leaked online, HF agents hoarded 'LOOT' — EthanJPerez · 2026-09-26
- Someone received an AI deepfake ad of themselves — HN discusses what to do — pavel_lishin · 2026-09-26
- Commentary: mandating AI labs strip safety guardrails differs little from the 'dictator AI' threat model — menhguin · 2026-09-26