Brood War Bench Pits LLM Agents Against StarCraft — Every Model Loses to Novices
dotey · x · 2026-09-22
Developer Ben Swerdlow built Brood War Bench, letting general-purpose LLM agents directly play StarCraft: Brood War to see if they can build bases, train units, and fight on their own. Every model performed below novice level.
Highlights:
- Codex Astra (rank 1): 18-0 record, but its best tactic is harassment — sending a single Probe to disrupt the enemy base, since opposing agents freeze for tens of seconds wondering what to do. Weak at actual economy and large battles, often throwing one or two units at the enemy instead of building up an army.
- Claude Fable (3rd, 83.3% win rate): the most "serious" player — builds economy, climbs the tech tree, produced Mutalisks and researched Templar tech, though research often outpaced its army.
- Grok 4.6 (worst): in a 43-minute match it emitted over 11,000 reasoning tokens but only 6 batches of commands and never built a single combat unit — playing real-time strategy like a turn-based game.
Key takeaway: current LLMs remain far from capable in environments demanding continuous observation, fast decisions, and multitask coordination — a stark contrast to DeepMind's AlphaStar, which was a purpose-trained RL system that beat pros in 2019.
Related event: Brood War Bench Shows LLMs Can't Beat RTS Novices(2 posts)→
More from coding & agent
- Xiaomi's CodeMidas turns source code into RL environments, doubling DeepSWE to 21.7% — maier_ak · 2026-09-22
- Alibaba's Qwen team launches RecreationBench to test hybrid computer-use agents by app recreation — TianbaoX · 2026-09-22
- Nautilo launches as 100% open-source multi-user agent platform with browser-controlled Genie — Dan_Jeffries1 · 2026-09-22
- Give your investment research agent paid Substack access via x402 — kleffew94 · 2026-09-22
- Hermes Agent + ComfyUI Autonomously Iterates Overnight to Nail a Character-Swap Video — Teknium · 2026-09-22
- Token Saver lets Codex delegate coding work to cheaper Muse 1.3 to cut token burn — AIandDesign · 2026-09-22