DiG-bench: Frontier AI Models Still Stumped by Simple Text-Based Discovery Games

jcrwhittington · x · 2026-08-12

A research team introduced DiG-bench, a new benchmark of 70 text-based interactive games designed to evaluate LLMs' scientific discovery and rule-exploration capabilities.

Results show that while frontier models have improved significantly in recent months, they are still stumped by surprisingly simple text-based discovery problems. A substantial bottleneck remains in native text-domain exploration.

Related event: DiG-bench: Frontier LLMs Still Stumble on Simple Text Discovery Games(12 posts)→

Original post →

More from Research

Research channel →