Tencent's GameHorizon: 5,000 Hours of AAA Gameplay Data to Benchmark 47 Models
tencent · hf · 2026-09-22
Tencent releases GameHorizon Suite, a unified data and evaluation suite for measuring AI gameplay ability across temporal horizons:
- GameHorizon-Annotator: a scalable pipeline for auto-annotating multi-horizon instructions
- GameHorizon-Data: the first large-scale AAA gameplay dataset — 5,000 hours from 21 games with temporally aligned video, player actions, and multi-horizon instructions, collected by 100 expert players
- GameHorizon-Bench: reproducible offline (thousands of standardized questions) and stepwise online tracks; the online track tests whether offline scores track real gameplay and localizes long-horizon failures
47 models were evaluated with over 1 million invocations, revealing a clear task-difficulty hierarchy and large capability gaps across model families. Dataset, annotator, and benchmark will be open-sourced.
More from Research
- Additive averaging kernels speed up finite Markov chains via partition optimization — michaelchchoi · 2026-09-22
- Non-asymptotic stability bounds for multivariate ensemble Kalman filters under Wishart fluctuations — michaelchchoi · 2026-09-22
- Randomized step sizes make Metropolis–Hastings robust to tuning, study finds — michaelchchoi · 2026-09-22
- Exact MCMC via Bernoulli factories for proposals with intractable normalizing constants — michaelchchoi · 2026-09-22
- ScienceBuddy-Jev answers plant biochem question in 0.62s at 99.97% confidence — Scobleizer · 2026-09-22
- Laya: 11.6k-star open-source engine outputs typed decisions in 33ms, no generation — pandeyparul · 2026-09-22