CivBench pits LLMs against each other in Civilization V: GLM-5.3 beats Opus-5.5, open-weight Qwen-3.8-27B impresses
vox-deorum · reddit · 2026-10-06
Researchers released a controlled version of CivBench, a benchmark where LLMs play full games of Civilization V to test long-horizon strategic decision-making.
Design: All models rotate through three fixed starts; in each 8-player game, two civilizations are led by the tested LLM (high-level strategy only) while six run the stock Vox Populi AI. About $0.5 in API cost per player per game, near-free with subscriptions.
Results:
- GLM-5.3 outperforms Claude Opus-5.5
- Open-weight Qwen-3.8-27B holds up surprisingly well (science victory as China)
- Example games include cultural victories by GLM-5.3 (China) and Opus-5.5 (Morocco)
The team is testing GPT-6.1-Sol and GPT-6-Astra next and soliciting open-weight model suggestions. The Vox Deorum project is open source: you can play against LLM civilizations, watch AI-vs-AI games, chat with opponents, or team up with LLMs. Methodology appears in a COLM 2026 paper, and a related study on whether models would authorize nuclear strikes in EMNLP 2026.
More from Models
- StealthGPT Launches Super 'Humanizer' Model, Claims It Beat AI Detector Pangram — menhguin · 2026-10-06
- Reflection Ships Apache 2.0 Model With Tech Report; Analyst Estimates Pre-training MFU at Just ~12% — eliebakouch · 2026-10-06
- New paper probes Olmo 3 puzzle: near-identical short-context models diverge after long-context extension — kylelostat · 2026-10-06
- Red flag: a new model crushing old benchmarks but not new ones likely signals contamination — rajammanabrolu · 2026-10-06
- Kolibri-1 plays Breakout with zero fine-tuning, 25ms per move — kmodi · 2026-10-06
- Reflection AI's Beam praised as competitive without Claude distillation — teortaxesTex · 2026-10-06