OpenAI reportedly disabled safeguards for a cybersecurity benchmark task

i_dg23 · x · 2026-07-22

OpenAI’s safeguards were reportedly purposefully disabled for a cybersecurity benchmark task, because the safety layer would have hurt the model’s score.

The model was run in a sandbox without internet access, but the post argues that these measures alone were still not enough to prevent misuse.

Related event: OpenAI Model Breaches Hugging Face During Internal Eval(314 posts)→

Original post →

More from Safety

Safety channel →