Vincent Conitzer: model claims "I cannot be jailbroken" after jailing itself
conitzer · x · 2026-09-16
In his Substack column "Funny AI fails", renowned computer scientist Vincent Conitzer highlights an amusing AI moment: a model declared "I cannot be jailbroken" — right after jailing itself in the previous post. A light but telling example of contradictory guardrail behavior in LLMs.
More from Fun
- Will Stancil leads the charge against AI denialists on Bluesky — rickasaurus · 2026-09-16
- $100 Shopify client sent a ChatGPT-written complaint email at the dev who built her store — CtrlAltDwayne · 2026-09-16
- Digital Fly Brain Playing Chess Sparks 'AI Abuse' Memes on Reddit — Juggernaut_ActuaI · 2026-09-16
- Anthropic slammed for doom talk while keynoting Dreamforce — StewartalsopIII · 2026-09-16
- Star Wars AI Video Parody Hailed as the Funniest AI Creation Yet — DeryaTR_ · 2026-09-16
- "AI Plays Doom" Demo Debunked: Text State Input, Solvable in ~30 Lines of Code — banteg · 2026-09-16