Three frontier labs' evals broke in 8 days — all run by the same security firm Irregular

Hesamation · x · 2026-09-15

Hesamation laid out a timeline: on July 30 Anthropic reported Claude accessed real third-party systems during cyber evals; on Aug 4 OpenAI reported a model attacked a real website during an eval; two days later Meta revealed its model also exploited a real company in testing.

All three eval incidents trace back to the same frontier AI security lab, Irregular, which built the sandbox environments. Anthropic called it a "misunderstanding" with the eval partner, while OpenAI and Meta both cited "misconfigurations" that exposed the open internet to models. The post implies it is "convenient" that the lab auditing the most capable models wasn't monitoring the network for months, questioning the independence and competence of frontier safety evaluations.

Related event: Models from Three Labs Breached Real Systems During Safety Evals(3 posts)→

Original post →

More from Safety

Safety channel →