Cowcraft MCP and WoWBench go live, testing LLM agents inside World of Warcraft
djcows · x · 2026-10-11
djcows launched Cowcraft MCP and WoWBench to benchmark general intelligence of models by pointing agents at a World of Warcraft server; runs can be spectated in-game and compared on WoWBench, using an MMO as a live agent evaluation arena.
Related event: WoWBench Lets AI Agents Compete Inside World of Warcraft(2 posts)→
More from coding & agent
- AI sped up coding but broke QA: 64% of defects caught before prod in January — alex_verem · 2026-10-11
- Atlassian CPO's zero-to-one playbook for PMs: vibe-code first, then hire engineers — lennysan · 2026-10-11
- Dev hand-made 2 SwiftUI text animations, had Claude generate 182 more, open-sources all 184 — amos_gyamfi · 2026-10-11
- Ben Hylak on building simulations for agent evals: replay traces, detect sim awareness — HamelHusain · 2026-10-11
- Turingo detects AI writing by replaying document revision history, not text predictions — sethlazar · 2026-10-11
- gemini-cli VSCode extension leaked disposables due to comma-expression bug in activate() — nosmile99 · 2026-10-11