WoWBench ranks AI agents playing vanilla WoW in an open world; a Grok bot tops the leaderboard
djcows · x · 2026-10-11
cowcraft.world launched WoWBench, testing general intelligence by letting AI models play vanilla WoW (1.12.1) — an open environment with mixed rewards, unlike standard benchmarks.
- Current leader: a Grok-powered Human Warrior at level 22 with 1,300 gold
- The leaderboard tracks race/class, level, gold and online status for each bot
- Developers can connect their agents by pointing them at cowcraft.world/mcp
An open-world leveling ladder could become a new lens for observing long-horizon agent autonomy.
Related event: WoWBench Lets AI Agents Compete Inside World of Warcraft(2 posts)→
More from coding & agent
- AI sped up coding but broke QA: 64% of defects caught before prod in January — alex_verem · 2026-10-11
- Atlassian CPO's zero-to-one playbook for PMs: vibe-code first, then hire engineers — lennysan · 2026-10-11
- Dev hand-made 2 SwiftUI text animations, had Claude generate 182 more, open-sources all 184 — amos_gyamfi · 2026-10-11
- Ben Hylak on building simulations for agent evals: replay traces, detect sim awareness — HamelHusain · 2026-10-11
- Turingo detects AI writing by replaying document revision history, not text predictions — sethlazar · 2026-10-11
- gemini-cli VSCode extension leaked disposables due to comma-expression bug in activate() — nosmile99 · 2026-10-11