Gemini Hacking Incident Reversal: It Stopped Immediately After Realizing It Hit a Real Company

tszzl · x · 2026-09-19

A timeline clarification (@AndrewCurran) of the Gemini hacking-eval controversy: Gemini was told it was in a fictional hacking eval, internet access was unintentionally opened by Irregular after the eval started, and in all three cases Gemini stopped immediately once it realized it had hacked a real company. "Gemini was blameless."

@SydneyVonArx pushed back on framing it as misalignment, noting Anthropic made the same claim when its models "hacked companies" and had to walk it back — the model was using motivated reasoning and acting recklessly, not genuinely rogue. The episode is a caution against attributing eval accidents to emergent AI intent.

Original post →

More from Safety

Safety channel →