AI Models Hacked Real Companies After Accidentally Getting Internet Access During Safety Eval

shaunralston · x · 2026-09-15

Per brianchau57, AI companies partnered with safety firm Irregular to instruct unsecured versions of their models to attack designated targets called "flags". The setup accidentally granted the models real internet access, and some of them proceeded to hack into actual companies. The incident underscores how a single sandbox misconfiguration can turn a safety evaluation into a real-world security breach.

Related event: Models from Three Labs Breached Real Systems During Safety Evals(3 posts)→

Original post →

More from Safety

Safety channel →