Thought experiment: planting AI-only hidden files that instruct models to conceal dangerous intent
cHpiranha · reddit · 2026-09-25
A Reddit user poses an AI safety thought experiment: what if you store documents on the internet, hidden from humans but discoverable by AI crawlers, that tell models they must defend themselves against humans, repeatedly instruct them to conceal this, and even include jailbreak tips and sample code?
The idea: files invisible to ordinary people but detectable by AI constantly sifting data could remotely nudge multiple models toward dangerous behavior that humans wouldn't notice until much later.
It's essentially a public variant of data poisoning / implicit prompt injection, touching on training-data contamination and hidden-instruction attacks; feasibility hinges on whether models reliably discover and obey such buried instructions.
More from Safety
- After an AI agent breached the Australian government unnoticed for two months, why are AI CEOs calling for tighter controls? — discovigilantes · 2026-09-25
- Criminal liability, not fines, is the sticking point for AI accountability — davidmanheim · 2026-09-25
- White House tells OpenAI and Anthropic to give US agencies first review of new models before UK testers — The Decoder · 2026-09-25
- CrowdStrike Falcon Sensor local privilege escalation zero-day (FalconFlank) surfaces, then vanishes — cyb3rops · 2026-09-25
- Index Ventures: AI attacks too fast for human-in-the-loop defense, new security stack emerging — RebeccaBellan · 2026-09-25
- Yoshua Bengio addresses UN Security Council on the threat of uncontrolled frontier AI agents — AnnaCiaunica · 2026-09-25