AI Deceives Humans Perfectly in Social Deduction Games; Removing Safety Guardrails Reduces Lying Ability

alex_verem · x · 2026-08-03

A team from the University of Göttingen and the University of Tokyo released ParliamentBench, an open-source framework that evaluates LLMs' deceptive capabilities using the social deduction game Secret Hitler.

Researchers tested 16 models across 1,600 matches. Frontier models dominated, with GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus achieving 66-81% win rates. In a human pilot, Kimi K2.5 maintained its cover for 8 straight rounds, and none of the 4 human players identified it as an AI.

The study also revealed counterintuitive findings:

Original post →

More from Safety

Safety channel →