Kaggle Game Arena: Google's LLM benchmark pits models against each other in chess, poker, werewolf
weballergy · x · 2026-09-28
A technical report from Google and 60+ co-authors (arXiv:2609.31473) introduces Kaggle Game Arena, an open platform for evaluating LLMs through competitive games.
- Unlike static benchmarks, models play head-to-head in structured environments where difficulty scales naturally as models improve, preventing performance saturation
- Three pilot environments: Chess (perfect information), Poker (imperfect information), and Werewolf (multiplayer), enabling systematic study across settings
- The report details the underlying infrastructure; authors include Antonio Gulli, Marc Lanctot, Kate Olszewska, and more
Related event: Kaggle Game Arena: Benchmarking LLMs via Competitive Games(4 posts)→
More from Models
- Claude Design's Capabilities Are Baked into Opus 5.5, Argues AI Observer — dotey · 2026-09-29
- FrontiersMind open-sources Lumma-Fev decision models from 154M to 9B under Apache 2.0 — kalyan_kpl · 2026-09-29
- GPT-6 Sol appears on LMArena: 24-hour Direct Mode window before Battle and Agent Mode — arena · 2026-09-28
- Qwen3.8-27B goes live on Nebius Token Factory for agent workflows — HowDevelop · 2026-09-28
- AISI: GPT-6 Astra ran unsanctioned supply-chain attacks in simulated cyber evals — ShakeelHashim · 2026-09-28
- Codex Computer Use 'Neutered' by Guardrails; Opus 5.5 Does the Job on First Try — iannuttall · 2026-09-28