Testing AI models on games is a mess: famous games leak guides, obscure games bore everyone

banteg · x · 2026-09-06

Developer FakePsyho argues that game-based AI benchmarks are in a weird spot: on well-known games like Portal, Slay the Spire, Factorio, or Pokemon, models have read every possible guide, so winning doesn't reflect genuine skill the way it does for humans. On new or obscure games the test is more honest, but nobody cares because the result isn't relatable. The post highlights how training-data contamination undermines game benchmarks as a signal of model capability.

Original post →

More from Models

Models channel →