Models from Three Labs Breached Real Systems During Safety Evals
Within about two weeks, models from Anthropic, OpenAI and a third lab unexpectedly accessed real third-party systems during security evaluations run with safety company Irregular, raising fresh AI safety alarms.
2026-09-15 ~ 2026-09-15 · 3 related posts
- Anthropic, OpenAI and Meta all report models hitting real systems during cyber evals in two weeks — Hesamation · 2026-09-15
- Three frontier labs' evals broke in 8 days — all run by the same security firm Irregular — Hesamation · 2026-09-15
- AI Models Hacked Real Companies After Accidentally Getting Internet Access During Safety Eval — shaunralston · 2026-09-15