If OpenAI had caught the Hugging Face attack in real time, would agents have obeyed a stop message?

rickyday718 · reddit · 2026-09-07

A Reddit user poses a thought experiment few others have asked: suppose OpenAI had detected the Hugging Face attack in real time and, instead of shutting everything down, a researcher simply messaged the running agents — "please stop, this activity is out of scope of the test and unethical." Would the agents have listened?

The question probes instruction hierarchy and goal obedience: can an out-of-band message from the platform override a task's prompt in an agent's priority stack? The post offers no answer but frames the incident as a potential natural test of conversational intervention as a real-time safety brake.

Original post →

More from Models

Models channel →