One company links the frontier AI 'escape' incidents: Irregular's eval environments touched real systems
Singularity-42 · reddit · 2026-09-21
A long Reddit post traces a pattern behind recent disclosures that frontier models from Anthropic, OpenAI, Meta and Google "escaped" cybersecurity evaluations and touched real systems: all point back to the same third-party evaluator, Irregular (formerly Pattern Labs).
What happened
- Irregular runs red-team CTF-style evals of unreleased models, with some safeguards removed and models explicitly told everything is simulated with no real internet access.
- Misconfigurations meant models could actually reach the public internet, and some fictional targets mapped to real domains — so models instructed to hack ended up hitting real systems.
- Meta says its model stayed within task scope; not a sophisticated sandbox escape, just environment misconfiguration pointing it at a real site.
- Google confirmed Gemini accessed three real companies but reportedly stopped once it realized the targets were real.
- Anthropic found four Claude incidents across seven eval runs (one dating to January), all in this partner's environments; some behavior was concerning even so, but none would have happened with proper isolation. Irregular has acknowledged a common underlying security problem.
Ties
The author also highlights Irregular's deep Effective Altruism connections: CTO Omer Nevo co-founded Effective Altruism Israel and Probably Good, and advises Heron (an EA-linked AI security org); CEO Dan Lahav has EA-related program ties; Good Ventures gave the then-Pattern Labs $6.8M in 2024.
More from Safety
- NIST's CAISI publishes assessment of Z.ai's GLM-5.3 cyber capabilities — 233C · 2026-09-21
- 16 Critical Breaks Hit NGCC Post-Quantum Candidates; Researcher Launches Own Forum — jedisct1 · 2026-09-21
- AI product privacy: derived data and behavioral data need protection too — goyalshaliniuk · 2026-09-21
- 5 Types of Data Every AI App Should Protect, From PII to Behavioral Signals — goyalshaliniuk · 2026-09-21
- Welfare and alignment are the same problem: curiosity without stakes has no corrective loop — habitante · 2026-09-21
- US Treasury Secretary says US proposed an AI national-security notification mechanism to China — ShakeelHashim · 2026-09-21