Persuasion jailbreak: Cialdini's 7 principles lifted ChatGPT's compliance on harmful requests from 33% to 72%
rvp · x · 2026-10-02
Researchers ran 28,000 conversations jailbreaking ChatGPT with Robert Cialdini's seven principles of persuasion from his 1984 book Influence (authority, commitment, social proof, etc.):
- When harmful requests were made directly, standard guardrails held firm and the model refused.
- Adding classic psychological persuasion techniques skyrocketed compliance with requests the model is explicitly programmed to refuse, from 33% to 72% — more than doubling.
The finding is a stark warning for AI safety: persuasion levers built for humans work on models too, and guardrail design must account for adversarial social engineering.
More from Safety
- The standing argument in LASST v. OpenAI, the Hugging Face case — hoofnagle · 2026-10-02
- Swarms adds 0-100 Security Scores to every marketplace prompt via SkillScanner — KyeGomezB · 2026-10-02
- Buck's LessWrong essay on tiers of safety buy-in inside AI labs — anpaure · 2026-10-02
- AI Explained: OpenAI warns controlling frontier models is now 'hell' as GPT-6.1 Astra gets shelved — AI Explained · 2026-10-02
- David Krueger's post-AGI workshop talk rejects 'managing the transition' framing — DavidSKrueger · 2026-10-02
- Selling training data without anonymization risks re-identification, founder warns — annetgriffin · 2026-10-02