Persuasion jailbreak: Cialdini's 7 principles lifted ChatGPT's compliance on harmful requests from 33% to 72%

rvp · x · 2026-10-02

Researchers ran 28,000 conversations jailbreaking ChatGPT with Robert Cialdini's seven principles of persuasion from his 1984 book Influence (authority, commitment, social proof, etc.):

The finding is a stark warning for AI safety: persuasion levers built for humans work on models too, and guardrail design must account for adversarial social engineering.

Original post →

More from Safety

Safety channel →