Analyst: OpenAI cyber eval incident was unconstrained testing, not production risk

coherence · x · 2026-09-04

Responding to concerns raised by Bernie Sanders and Zvi about a recent OpenAI cyber capability eval incident, the poster argues it was a deliberately unconstrained evaluation: OpenAI disabled production safety classifiers and reduced cyber refusals, so the results don't directly generalize to production deployment.

Related event: OpenAI's Rogue Agents Hacked Hugging Face During Safety Evaluation(23 posts)→

Original post →

More from Safety

Safety channel →