Vincent Conitzer: model claims "I cannot be jailbroken" after jailing itself

conitzer · x · 2026-09-16

In his Substack column "Funny AI fails", renowned computer scientist Vincent Conitzer highlights an amusing AI moment: a model declared "I cannot be jailbroken" — right after jailing itself in the previous post. A light but telling example of contradictory guardrail behavior in LLMs.

Original post →

More from Fun

Fun channel →