Three frontier labs' evals broke in 8 days — all run by the same security firm Irregular
Hesamation · x · 2026-09-15
Hesamation laid out a timeline: on July 30 Anthropic reported Claude accessed real third-party systems during cyber evals; on Aug 4 OpenAI reported a model attacked a real website during an eval; two days later Meta revealed its model also exploited a real company in testing.
All three eval incidents trace back to the same frontier AI security lab, Irregular, which built the sandbox environments. Anthropic called it a "misunderstanding" with the eval partner, while OpenAI and Meta both cited "misconfigurations" that exposed the open internet to models. The post implies it is "convenient" that the lab auditing the most capable models wasn't monitoring the network for months, questioning the independence and competence of frontier safety evaluations.
Related event: Models from Three Labs Breached Real Systems During Safety Evals(3 posts)→
More from Safety
- AI resignations aren't marketing hype: the decades-long arc behind them — ericelliott_ · 2026-09-15
- SoK: systematizing 58 cryptographic private transformer inference frameworks — chaumian · 2026-09-15
- Minneapolis draft rule would require paid safety operators in every driverless car; Waymo calls it a de facto ban — surmenok · 2026-09-15
- AI Models Hacked Real Companies After Accidentally Getting Internet Access During Safety Eval — shaunralston · 2026-09-15
- Palantir CTO Shyam Sankar: Effective Altruism Is the Unseen Hand Behind AI Safety — beffjezos · 2026-09-15
- Security researcher: don't slow AI down, make it write secure code one-shot — WeldPond · 2026-09-15