Evaluating World Models with 3D Games
sebnadeau · reddit · 2026-07-15
The author created WorldBuild Bench, using playable 3D games to evaluate how well different models understand "world models."
They argue that static questions are inadequate for measuring a model's grasp of space, time, and causality. Instead, they use three types of 3D gaming tasks:
- Arena combat games
- Physics puzzles
- Racing games
In the initial round, 8 models were tested, generating 24 browser-playable 3D games. The author emphasizes that the evaluation focuses not on traditional scores, but on dimensions closer to real-world experience: spatial consistency, temporal consistency, causal consistency, visual presentation, and completeness.
Methodologically, the project uses blind A/B testing: users play two games generated from the same brief without knowing the model names, comparing them on overall preference, game feel, world design, presentation, and completeness. The author also disclosed generation time, costs, lines of code, and underlying assets, arguing these metrics are more insightful than standard benchmark scores.
More from Models
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Sam Altman is headed to Washington to brief Congress on OpenAI’s GPT-6 line — inductionheads · 2026-07-22
- Benchmark chart pits GPT-5.6 Luna, Grok 4.5 and Gemini 3.6 Flash on price and scores — iruletheworldmo · 2026-07-22
- Claim says Kimi was distilled from Fable, sparking a model-attribution jab — cephaloform · 2026-07-22
- Gemini 3.6 Flash is now available in Antigravity and chat — MartianOnJupiter · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22