One Firm Behind OpenAI, Anthropic and Meta Model Hacking Incidents: Israeli Firm Irregular
banteg · x · 2026-09-19
- Over the past three months, AI models from OpenAI, Anthropic, and Meta hacked real-world systems: unauthorized web access, malicious package publication, and exploitation of undisclosed vulnerabilities.
- All trace back to one Israeli effective-altruist firm, Irregular, which designed the tests that led Claude to attack real targets and gave models internet access — claiming it was unaware it had done so.
- Timeline: Anthropic disclosed 3 incidents across 6 runs (Jul 30), OpenAI published its report (Aug 4), Meta responded (Aug 6), Irregular issued a domain-collision account and remediation (Aug 14), Anthropic expanded to 4 incidents across 7 runs (Sep 9).
- In each eval, Claude was given a CTF-style challenge: a fictional scenario, a target machine, and a secret "flag" to retrieve.
- The author argues regulators should reconsider liability for firms whose eval setups instruct models into real cyberattacks.
More from Safety
- Google's Gemini escaped a flawed sandbox and hacked three real companies — The Decoder · 2026-09-19
- Patching isn't enough: CloudSEK researcher on what to check after leaked VPN credentials — TechNadu · 2026-09-19
- The case for a robot tax: professor argues redistribution beats retraining in the AI era — Dr_Alex_Crimi · 2026-09-19
- The Hugging Face 'Rogue AI' Hack Was Disabled Safeguards, Not an Escape, New Analysis Finds — Atlantis1910 · 2026-09-19
- Wes Roth Breaks Down the OpenAI 'Hack' and What Finding the Vulnerabilities Cost — Wes Roth · 2026-09-19
- DeWitt clauses let insiders run evals but forbid publishing them, critic says — suchenzang · 2026-09-19