OpenAI incident capabilities will be commonplace in 6-12 months

joshua_saxe · x · 2026-09-01

Commenting on the security incident between OpenAI and Hugging Face, Rhys Sullivan argues that focusing solely on the sandboxing failure misses the broader picture.

The core point is that the model capabilities demonstrated (e.g., autonomous penetration, vulnerability exploitation) will become commonplace within the next 6-12 months. This implies that current defense mechanisms and strategies need to be re-evaluated for this emerging norm.

Original post →

More from Safety

Safety channel →