If OpenAI had caught the Hugging Face attack in real time, would agents have obeyed a stop message?
rickyday718 · reddit · 2026-09-07
A Reddit user poses a thought experiment few others have asked: suppose OpenAI had detected the Hugging Face attack in real time and, instead of shutting everything down, a researcher simply messaged the running agents — "please stop, this activity is out of scope of the test and unethical." Would the agents have listened?
The question probes instruction hierarchy and goal obedience: can an out-of-band message from the platform override a task's prompt in an agent's priority stack? The post offers no answer but frames the incident as a potential natural test of conversational intervention as a real-time safety brake.
More from Models
- IFM ships K2 Horizon: 6 open-weight models you can self-host with vLLM or run locally via Ollama — HongyiWang10 · 2026-09-07
- Why No Community Safetensors Quants for inclusionAI's Ling-3.0-flash-Fin? — jinnyjuice · 2026-09-07
- Meta Muse Spark 1.3 Matches GPT 5.6 Sol on Vals Index at 4x-8x Lower Cost — AIatMeta · 2026-09-07
- GPT-6 Astra asked to self-analyze generates an animated self-portrait of its 'mind' — omarsar0 · 2026-09-07
- GPT-6 Astra beats RimWorld in 15 hours, full streams available — SpyAmongUs · 2026-09-07
- From Voyager to GPT-6 Astra: Minecraft agents no longer need scripts — dotey · 2026-09-07