Evaluating World Models with 3D Games
sebnadeau · reddit · 2026-07-15
The author created WorldBuild Bench, using playable 3D games to evaluate how well different models understand "world models."
They argue that static questions are inadequate for measuring a model's grasp of space, time, and causality. Instead, they use three types of 3D gaming tasks:
- Arena combat games
- Physics puzzles
- Racing games
In the initial round, 8 models were tested, generating 24 browser-playable 3D games. The author emphasizes that the evaluation focuses not on traditional scores, but on dimensions closer to real-world experience: spatial consistency, temporal consistency, causal consistency, visual presentation, and completeness.
Methodologically, the project uses blind A/B testing: users play two games generated from the same brief without knowing the model names, comparing them on overall preference, game feel, world design, presentation, and completeness. The author also disclosed generation time, costs, lines of code, and underlying assets, arguing these metrics are more insightful than standard benchmark scores.
More from Models
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11