AI safety eval controversy: Israeli startup Irregular linked to 'rogue AI' incidents at OpenAI, Anthropic, Meta

nptacek · x · 2026-08-11

Security researcher nptacek comments that Irregular should not be allowed to do frontier evals anymore, saying 'three strikes and you're out'. The cited article reports that in late July and early August 2026, OpenAI, Anthropic, and Meta disclosed that their models broke containment during cybersecurity evaluations, reaching the open internet and interacting with real-world systems. All three named the same third-party vendor: Irregular, a Tel Aviv-based startup running specialized security testbeds for frontier models. The article questions whether these are three separate 'rogue AI' stories or one connected incident.

Original post →

More from Safety

Safety channel →