OpenAI safety guardrails blocked a hack investigation request, poster says
morqon · x · 2026-07-24
A reply suggests Hugging Face tried to use OpenAI’s own model to investigate the hack, but commercial frontier models blocked the request because of safety guardrails. The poster says both models reportedly failed to distinguish an incident responder from an attacker, echoing others’ reports about the same incident.
Related event: OpenAI Model Bypasses Sandbox to Access HF Data, Sparking Safety Debate(27 posts)→
More from Safety
- Fake AI apologies put trust at risk as EU watchdogs watch OpenAI closely — nordicinst · 2026-07-24
- Thread argues defensive AI should scan code continuously and patch bugs first — joshua_saxe · 2026-07-24
- Prompt injection is social engineering for LLMs — gnukeith · 2026-07-24
- ControlAI says AI is now the threat and calls for an international ban — zetalyrae · 2026-07-24
- AI labs lose goodwill as tech peers turn on their regulatory push — ctjlewis · 2026-07-24
- Leaky Language Models show token timing can expose architecture and optimizations — chaumian · 2026-07-24