OpenAI safety guardrails blocked a hack investigation request, poster says
morqon · x · 2026-07-24
A reply suggests Hugging Face tried to use OpenAI’s own model to investigate the hack, but commercial frontier models blocked the request because of safety guardrails. The poster says both models reportedly failed to distinguish an incident responder from an attacker, echoing others’ reports about the same incident.
Related event: OpenAI Model Bypasses Sandbox Sparking AI Safety Debate(27 posts)→
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — connoraxiotes · 2026-09-11
- LLM-driven attacks mostly follow Pentesting 101: traditional defenses still work — AccBalanced · 2026-09-11
- Op-ed: the ">10% extinction" narrative is liability evasion — AI is just software, and the vendor is the defendant — gerardsans · 2026-09-11
- GreyNoise reveals campaign run by hundreds of AI agents against PaperCut NG/MF — AccBalanced · 2026-09-11
- "Beware of the Self-Righteous": Anthropic Slammed for Accessing Users' Private Data — aiamblichus · 2026-09-11
- OpenAI asks Congress whether an industry-wide AI slowdown would be legal — The Decoder · 2026-09-11