Frontier Models Stumped by Simple Games: DiG-bench Tests AI Discovery

jcrwhittington · x · 2026-08-12

Researchers introduced DiG-bench, a new benchmark for evaluating the scientific discovery capabilities of LLMs. It features 70 text-based interactive games where models must experiment to uncover hidden rules and solve challenges.

By using purely text-based environments, the benchmark isolates discovery skills without visual confounds. Results show that while frontier models have improved, they still struggle with surprisingly simple problems that humans can beat on their first attempt.

Related event: DiG-bench: 70 Text Games Expose Reasoning Gaps in Frontier Models(13 posts)→

Original post →

More from Research

Research channel →