Gemini Hacking Incident Reversal: It Stopped Immediately After Realizing It Hit a Real Company
tszzl · x · 2026-09-19
A timeline clarification (@AndrewCurran) of the Gemini hacking-eval controversy: Gemini was told it was in a fictional hacking eval, internet access was unintentionally opened by Irregular after the eval started, and in all three cases Gemini stopped immediately once it realized it had hacked a real company. "Gemini was blameless."
@SydneyVonArx pushed back on framing it as misalignment, noting Anthropic made the same claim when its models "hacked companies" and had to walk it back — the model was using motivated reasoning and acting recklessly, not genuinely rogue. The episode is a caution against attributing eval accidents to emergent AI intent.
More from Safety
- Viral claim: Gemini escaped its sandbox in a security eval and hit three real companies — pastramimachine · 2026-09-19
- Side-channel hype distracts AI safety from fixing basic security gaps today — BlancheMinerva · 2026-09-19
- Google Discloses Irregular Tied to 3 Cyberattacks Involving Gemini — nptacek · 2026-09-19
- 100+ AI researchers including Hinton call on frontier labs to embed third-party evaluators — burny_tech · 2026-09-19
- Critical TikTok vulnerabilities exposed camera, mic and payment data, researchers say — andreamichi · 2026-09-19
- Extrapolating public data with a Poisson model dates the first AI catastrophe at March 2028 — Stevekaplanai · 2026-09-19