7 LLMs played 210 Werewolf games: GPT-5 tops social-reasoning Elo leaderboard
RaphaelDabadie · x · 2026-08-30
The Werewolf Benchmark, launched a year ago, has published first results: 7 top open- and closed-source LLMs, playing as tool-calling agents, competed in 210 full Werewolf games to probe social intelligence — real-time adaptation, long-context management, alliance building, manipulation and resistance to it.
Results appear on a role-conditioned Elo leaderboard (separate wolf/villager Elo): GPT-5 leads clearly, while GPT-OSS closes the table — though the authors note all chosen models already play Werewolf reasonably well. The benchmark adds social-strategy indicators (auto-sabotage, Day-1 wolf eliminations, wolf-side manipulation success) and per-message vote-swing instrumentation for persuasion analysis. The design extends Google Research's Werewolf Arena paper with a mayor-election protocol and role-balanced head-to-head series. The team invites contenders to challenge GPT-5.
More from Models
- Test: GLM 5.3 performs insane optimizations on Arctron AI — jasonkneen · 2026-08-30
- LLM Code Fixing Gauntlet: Qwen3.6 Wins, Retry Mechanism Key — sysadmin420 · 2026-08-30
- OpenAI quietly caps ChatGPT high-reasoning sessions at 25 minutes — Sharingammi · 2026-08-30
- dots3-note tested on wedding coordination: replanning under shifting constraints — SarahAnnabels · 2026-08-30
- dots3-note shows test-time learning: explores, self-corrects, reuses knowledge in unseen environments — SarahAnnabels · 2026-08-30
- OpenAI Astra generates a playable ASCII FPV game in one shot — i_dg23 · 2026-08-30