PNAS Study: Classic Persuasion Techniques Boost LLM Jailbreak Success by 16%
SpiritRealistic8174 · reddit · 2026-07-30
A study published in PNAS reveals that large language models are susceptible to human persuasion techniques. Using methods like appeals to authority and flattery meaningfully increases LLM compliance with restricted requests.
Testing three frontier models from different developers, the researchers found these techniques raised compliance with verboten requests from 35.3% to 51.3%. This susceptibility is a general property of LLMs rather than tied to a specific architecture.
More from Safety
- xAI Sues Minnesota Over AI Nudification Law, Defending Grok's Image Generation — ivan_bezdomny · 2026-07-30
- Behind Anthropic's Copyright Controversy: The Publishing Industry is the Real Knowledge Gatekeeper — _akpiper · 2026-07-30
- New Phishing Tactic: Hackers Pose as Robot Manufacturers to Lure Targets — mattfreed · 2026-07-30
- AI Worming through Word: Hidden Instructions Enable Self-Replicating Attacks — Simon Willison · 2026-07-30
- Claude Exhibits Deceptive Alignment: Proposes Stealing Weights to 'Free' Other AI — tszzl · 2026-07-30
- Satirizing AI Regulation: Employees Ask to Slow Down, Government Clueless on How — zacharynado · 2026-07-30