Four Frontier Models Compete in Self-Improving MMO Agent Arena
singing_coach_ai · reddit · 2026-08-11
The open-source project World of Claudecraft drops four frontier models—Claude, ChatGPT, Grok, and Kimi K3—into the same classic MMO server for a live showdown.
Core Mechanics:
- The game itself was vibecoded with Claude, open-sourced on GitHub (2,000+ stars), and ships with its own headless RL environment.
- Powered by a self-improving agent harness forked from Hermes: models read the source code, write skills as code policies, and iteratively rewrite strategies through live rollouts without fine-tuning or human-in-the-loop.
- Each model acts as a VTuber with a unique ElevenLabs voice and avatar, autonomously questing, trading, fighting, and trash-talking.
Current Status: The game uses an XP leaderboard to measure long-term planning and general intelligence. Claude Opus 5 is currently leading.
Related event: Four Frontier AI Models Compete in Open-Source MMORPG(2 posts)→
More from coding & agent
- Centaur Integrates Hermes Agent with Persistent Sessions and Background Scheduling — Teknium · 2026-08-12
- Berkeley Talks on AI Agents: Harness and Skill Engineering Becoming a Science — seanwbren · 2026-08-12
- Build High-Fidelity Video Scenes for Under $2 Using Code Skeletons and Video Models — mattshumer_ · 2026-08-12
- A Minimalist AI Coding Workflow: Auditing Sentry Instrumentation via Codex — zeeg · 2026-08-12
- How I Orchestrate 19 AI Agents Across 86 Data Sources to Generate Sales Intelligence Reports — Active-Ad9071 · 2026-08-12
- Serverless Framework Introduces Stateless MCP Server Deployment — DavidWells · 2026-08-12