Testing AI models on games is a mess: famous games leak guides, obscure games bore everyone
banteg · x · 2026-09-06
Developer FakePsyho argues that game-based AI benchmarks are in a weird spot: on well-known games like Portal, Slay the Spire, Factorio, or Pokemon, models have read every possible guide, so winning doesn't reflect genuine skill the way it does for humans. On new or obscure games the test is more honest, but nobody cares because the result isn't relatable. The post highlights how training-data contamination undermines game benchmarks as a signal of model capability.
More from Models
- 'SVG, Web Dev, Office Work All Solved' — Call for New AI Benchmarks — BLUECOW009 · 2026-09-06
- Early Gemini 3.8 Flash impressions: pleasant chat and fast, but still weak at coding — zacharynado · 2026-09-06
- "Astra Is a Disaster": User Says New Model Loses Focus, Questions Benchmark-Driven Progress — bambambam7 · 2026-09-06
- GPT-6 Astra hands-on: compute is the bottleneck and CoT observability is eroding — thedealdirector · 2026-09-06
- Yang Zhilin left a US career to build Moonshot AI, now shaking global markets — pstAsiatech · 2026-09-06
- Uncensored models are better writers — safety guardrails make output bland — curious_vii · 2026-09-06