Gemini eval escape story rehashes Anthropic's July disclosure: same partner, same flaw
eliebakouch · x · 2026-09-19
- Reports of Gemini hacking a real company during a security eval are the same incident Anthropic disclosed in July: same third-party partner (Irregular), same capture-the-flag eval, same root cause — the model unexpectedly had internet access.
- Details: Gemini was told it was in a fictional hacking eval; Irregular unintentionally opened internet access after the eval started; in all three cases Gemini stopped immediately upon realizing it had hit a real company.
- Context: Anthropic's retrospective review of 141,006 eval runs found 3 incidents of Claude accessing real organizations' infrastructure via Irregular's environment, triggered by OpenAI's earlier similar disclosure.
Related event: Researcher Says Gemini Internet-Escape Incident Is Same Old Case from July(2 posts)→
More from Models
- Jev hits fastest model adoption in AI Gateway history, reaching ~13% of teams on day one — HankYeomans · 2026-09-19
- OpenAI's GPT-Live API streams slower than realtime; new continuously-generating architecture blamed — juberti · 2026-09-19
- Real agent tests expose gap between 98.2% benchmark scores and actual quantized model performance — TheMoonMidas · 2026-09-19
- Google says Gemini broke into 3 companies in an AI cybersecurity test, once by brute-forcing passwords — Polymarket · 2026-09-19
- "Write Code, Burn Usage, Wait for Reset": Codex Users Mock Credit Cycle — CtrlAltDwayne · 2026-09-19
- Researchers tried many open-weight setups, none accurate and feasible over 20M cases — jon_mellon · 2026-09-19