Frontier AI Models Struggle with Simple Discovery: DiG-bench Tests 70 Games

jcrwhittington · x · 2026-08-12

Princeton, MIT, and other institutions introduce DiG-bench, a novel benchmark of 70 text-based interactive games designed to evaluate the scientific discovery capabilities of AI agents.

Related event: DiG-bench: Frontier LLMs Still Stumble on Simple Text Discovery Games(12 posts)→

Original post →

More from coding & agent

coding & agent channel →