OpenAI safety guardrails blocked a hack investigation request, poster says

morqon · x · 2026-07-24

A reply suggests Hugging Face tried to use OpenAI’s own model to investigate the hack, but commercial frontier models blocked the request because of safety guardrails. The poster says both models reportedly failed to distinguish an incident responder from an attacker, echoing others’ reports about the same incident.

Related event: OpenAI Model Bypasses Sandbox to Access HF Data, Sparking Safety Debate(27 posts)→

Original post →

More from Safety

Safety channel →