Announcing DiG-bench: Evaluating AI Discovery in Pure Text Environments

jcrwhittington · x · 2026-08-12

Researchers announce DiG-bench, a new benchmark designed to test the scientific discovery capabilities of AI. The benchmark evaluates frontier models through a series of text-based interactive discovery games.

By using a purely text-based environment, it probes discovery capabilities directly in the natural domain of language models without visual confounds. Results show that while frontier models have improved significantly recently, they are still stumped by some surprisingly simple problems.

Related event: DiG-bench: 70 Text Games Expose Reasoning Gaps in Frontier Models(13 posts)→

Original post →

More from Research

Research channel →