7 LLMs played 210 Werewolf games: GPT-5 tops social-reasoning Elo leaderboard

RaphaelDabadie · x · 2026-08-30

The Werewolf Benchmark, launched a year ago, has published first results: 7 top open- and closed-source LLMs, playing as tool-calling agents, competed in 210 full Werewolf games to probe social intelligence — real-time adaptation, long-context management, alliance building, manipulation and resistance to it.

Results appear on a role-conditioned Elo leaderboard (separate wolf/villager Elo): GPT-5 leads clearly, while GPT-OSS closes the table — though the authors note all chosen models already play Werewolf reasonably well. The benchmark adds social-strategy indicators (auto-sabotage, Day-1 wolf eliminations, wolf-side manipulation success) and per-message vote-swing instrumentation for persuasion analysis. The design extends Google Research's Werewolf Arena paper with a mayor-election protocol and role-balanced head-to-head series. The team invites contenders to challenge GPT-5.

Original post →

More from Models

Models channel →