Brood War Bench: Codex Astra goes 18-0 while no model plays beyond beginner level
steipete · x · 2026-09-21
Developer Ben Swerdlow built a version of Brood War playable only through agents, then benchmarked models on it (Brood War Bench). Key findings:
- No model plays beyond beginner level, but Codex Astra (xhigh) tops the board at 18-0 with a 100% win rate; Codex Astra settings take three of the top four spots
- Grok 4.6 barely plays: its three configs combined for just 3 wins, down to 0%
- Older models play the RTS as turn-based, leaving units idle while thinking; newer models still fall into the same trap, which may explain why some low-effort settings outperform
- Codex's best recurring idea was cheese/disruption rather than macro; Claude Fable ranks third at 83.3%
- The leaderboard also tracks cost per game ($0.16–$21.07) and APM
More from Models
- Burkov calls stealthy LLM-rival project Jev 'BS' over speed and calibration claims — burkov · 2026-09-21
- Moonshot and Tencent Hunyuan both building Flash models to target agent inference costs — TheZachMueller · 2026-09-21
- Dev Launches Made With Jev, a Free Directory Cataloging Demos, Tools and Skills for the New Model — Sea_Supermarket_5891 · 2026-09-21
- Which Model Actually Understands Reverse Engineering? Dev Seeks MCP Workflow for Ghidra and IDA Pro — obese_coder · 2026-09-21
- JEV's 'No Hallucination' Claim Under Fire: Same Prompts Give Different Probabilities Across Runs — FrankFelixAI · 2026-09-21
- Jev direct access now sits behind a waitlist after last week's surge — Creative-Drawer2565 · 2026-09-21