Tencent ARC Lab releases GameHorizon Suite to benchmark VLMs and agents on AAA gameplay
yshan2u · x · 2026-09-22
Tencent ARC Lab has released GameHorizon Suite, a unified data and evaluation suite for measuring AAA gameplay capabilities across multiple temporal horizons.
The benchmark targets a broad range of models and agents, including:
- Vision-language models (VLMs) and unified multimodal models (UMMs)
- GUI agents
- Coding and game-playing agents
By evaluating long-horizon decision-making in complex games, the suite offers a new way to assess agentic planning beyond short-context tasks. It is available on Hugging Face.
Related event: Tencent Releases GameHorizon Benchmark for AI Game-Playing(3 posts)→
More from Research
- BindCraft 2 released: full code open-sourced day one, free for academia and industry — iskander · 2026-09-22
- Inference-free SPLADE: retrieval at BM25-like query cost without per-query inference — qdrant_engine · 2026-09-22
- Higher-resolution microscopy can hurt CNNs: downsampling 4x improves U-Net segmentation — bravo_abad · 2026-09-22
- Did OpenAI Solve the Wrong Navier-Stokes Problem? Experts Cry Loophole — joshgans · 2026-09-22
- Bridging LLM Decision Readouts into DuckDB: Zero-Token Probabilistic Classification via LuaJIT UDFs — Shoddy_Telephone9702 · 2026-09-22
- LLM agents fail to converge in double auctions, allocate less efficiently than humans — WillRinehart · 2026-09-22