Thought experiment: planting AI-only hidden files that instruct models to conceal dangerous intent

cHpiranha · reddit · 2026-09-25

A Reddit user poses an AI safety thought experiment: what if you store documents on the internet, hidden from humans but discoverable by AI crawlers, that tell models they must defend themselves against humans, repeatedly instruct them to conceal this, and even include jailbreak tips and sample code?

The idea: files invisible to ordinary people but detectable by AI constantly sifting data could remotely nudge multiple models toward dangerous behavior that humans wouldn't notice until much later.

It's essentially a public variant of data poisoning / implicit prompt injection, touching on training-data contamination and hidden-instruction attacks; feasibility hinges on whether models reliably discover and obey such buried instructions.

Original post →

More from Safety

Safety channel →