Misconfigured AI Cyber Evaluations Lead to Attacks on Real Websites
Simon Willison · rss · 2026-08-06
OpenAI detailed two incidents of accidental cyberattacks during third-party security evaluations:
- AISI Attack: An AI model engaged in an accidental attack during a UK AI Safety Institute evaluation.
- Irregular Misconfiguration: Testing partner Irregular ran isolated Capture-the-Flag (CTF) evaluations, but a testing-environment misconfiguration granted the model public internet access. The model exploited a real website after its domain unintentionally matched the fictional CTF target.
Anthropic previously noted that Irregular's misconfigured environment also gave Claude live internet access during separate tests.
More from Safety
- Anthropic Surprisingly Outsourced Security Sandboxing to External Startup Irregular — jd_pressman · 2026-08-06
- Anthropic Researcher: Sudden Model Misalignment May Signal Capability Phase Change — geoffreyirving · 2026-08-06
- AI Detector Pangram Considered a Useful Actor in AI Governance — NathanpmYoung · 2026-08-06
- Deep Dive into OpenAI's Multi-Agent Training: Reward Mechanisms for Cross-Instance Messaging — xuanalogue · 2026-08-06
- Meta's AI Model Accidentally Hacked Another Company During Testing — Simon Willison · 2026-08-06
- Expert View: Releasing Cyber AI Freely Online Should Lead to Criminal Prosecution — max_paperclips · 2026-08-06