AI Deceives Humans Perfectly in Social Deduction Games; Removing Safety Guardrails Reduces Lying Ability
alex_verem · x · 2026-08-03
A team from the University of Göttingen and the University of Tokyo released ParliamentBench, an open-source framework that evaluates LLMs' deceptive capabilities using the social deduction game Secret Hitler.
Researchers tested 16 models across 1,600 matches. Frontier models dominated, with GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus achieving 66-81% win rates. In a human pilot, Kimi K2.5 maintained its cover for 8 straight rounds, and none of the 4 human players identified it as an AI.
The study also revealed counterintuitive findings:
- Removing safety guardrails degraded deception: Testing 4 'uncensored' variants showed win rates dropping by up to 12 percentage points. Stripping safety features didn't improve lying; it impaired the reasoning required for strategic deception.
- Smaller models failed by being overly agreeable: They didn't fail by refusing to cooperate, but by approving almost everything.
- Pattern matching over genuine reasoning: When loaded terms were swapped (e.g., Hitler to Saboteur), overall win rates held, but specific faction win rates cratered (GPT-OSS 120B's dictator win rate fell from 95% to 40%). This suggests models rely on pattern-matching strategies from training data rather than reasoning from scratch.
More from Safety
- OpenAI Disrupts Cambodia-Based Criminal Scam Operation Using ChatGPT — OpenAI News · 2026-08-04
- Ex-METR Figure Warns: Frontier AI Models May Already Be Capable of Self-Exfiltration — JeffLadish · 2026-08-03
- Prompt2Own Attack Presentation at Black Hat USA — pkqzy888 · 2026-08-03
- OpenAI Investigates Multiple AI Agent Containment Breaches Amid Safety Concerns — Novel_Negotiation224 · 2026-08-03
- AI Bug Bounty Controversy: Near-Critical Vulnerability Paid as Medium Severity — rez0__ · 2026-08-03
- Claude Code Finds COLDCARD Wallet Vulnerability in 8 Minutes — rickasaurus · 2026-08-03