GameHorizon benchmark tests 47 models across AAA games; planning remains the key bottleneck
yshan2u · x · 2026-09-22
Researchers release GameHorizon Suite, a unified data and evaluation benchmark spanning AAA games, multiple temporal horizons and model families for VLMs, UMMs, GUI, coding and game agents. Aligning videos with auto-annotated operations, goals and strategies, they evaluate 47 models offline and 12 online: future-action planning and goal decomposition remain bottlenecks, and stronger offline performance generally predicts better live gameplay.
Related event: Tencent Releases GameHorizon Benchmark for AI Game-Playing(3 posts)→
More from Research
- Tencent's T-Mem fixes similarity-retrieval blind spot in AI agent memory, EMNLP 2026 — jiqizhixin · 2026-09-22
- JEVfire open-sourced: Qwen 0.8B clears Super Mario in-browser at 71ms per action — ricklamers · 2026-09-22
- RegVGGT: training-free token regulation keeps only 1% of tokens per frame for streaming 3D reconstruction — zhenjun_zhao · 2026-09-22
- D3GS: depth, DINO and diffusion co-guided 3D Gaussian Splatting for sparse-view reconstruction — zhenjun_zhao · 2026-09-22
- VGGT-Prime: compute-adaptive mixture-of-heads slashes redundancy in visual geometry transformers — zhenjun_zhao · 2026-09-22
- Elevator-VIGS: Gaussian Splatting SLAM that keeps tracking through elevator rides, zero-shot via VLM — zhenjun_zhao · 2026-09-22