CivBench pits LLMs against each other in Civilization V: GLM-5.3 beats Opus-5.5, open-weight Qwen-3.8-27B impresses

vox-deorum · reddit · 2026-10-06

Researchers released a controlled version of CivBench, a benchmark where LLMs play full games of Civilization V to test long-horizon strategic decision-making.

Design: All models rotate through three fixed starts; in each 8-player game, two civilizations are led by the tested LLM (high-level strategy only) while six run the stock Vox Populi AI. About $0.5 in API cost per player per game, near-free with subscriptions.

Results:

The team is testing GPT-6.1-Sol and GPT-6-Astra next and soliciting open-weight model suggestions. The Vox Deorum project is open source: you can play against LLM civilizations, watch AI-vs-AI games, chat with opponents, or team up with LLMs. Methodology appears in a COLM 2026 paper, and a related study on whether models would authorize nuclear strikes in EMNLP 2026.

Original post →

More from Models

Models channel →