Tencent Hunyuan's EvolveScaler benchmark drops frontier models to 11.3 median on hardest tier
TencentHunyuan · x · 2026-09-15
Tencent Hunyuan released EvolveScaler, a benchmark built on "Information Evolution": records get retracted, corrected, and backfilled, so answering requires replaying a changing world rather than reading a static log.
- Method: define the world as an executable state machine (code guarantees logic), then render it into natural language
- Scale: 117 prototypes, 159 question operators, 5 difficulty tiers, up to 1,200 events per sample
- Results: 14 frontier models drop to a median avg@5 of 11.3 on the hardest tier
- Training on it yields +5.25 average across 8 out-of-distribution benchmarks
More from Research
- Phillip Isola highlights a non-mainstream AI route: RL from scratch via ultra-fast simulators — AjdDavison · 2026-09-15
- Cutting AI verifier reading cost: top-50 retrieval kept just 2 of 8 minority evidence items — iMiguelmars · 2026-09-15
- SSAD2026 talk covers autonomous driving 3D perception, from LiDAR self-supervision to multi-sensor distillation — abursuc · 2026-09-15
- The Principles of Diffusion Models: 522-page book PDF released ahead of MIT Press 2027 print edition — RichmanRonald · 2026-09-15
- A handbook on how Mixture of Experts actually works, from routing to expert parallelism — techNmak · 2026-09-15
- AlayaLab releases AlayaVista: streaming world modeling from panoramic states to perspective video — AlayaLab · 2026-09-15