DiG-bench: Top Frontier Models Clear Only 20% on Simple Discovery Games

misovalko · x · 2026-08-15

Researchers from Princeton, MIT, and others released DiG-bench, a new benchmark designed to evaluate AI discovery capabilities. Using text-based discovery games to probe models in their native domain (avoiding confounding visual elements), the study reveals that while frontier models have improved significantly over the last few months, they remain stumped by surprisingly simple problems. Notably, humans can beat these games on the first try, whereas the best model clears only 20%. The research also found that advanced agentic harnesses provided zero lift over a basic setup.

Related event: DiG-bench: Frontier AI Models Win Only 20% of Simple Discovery Games(2 posts)→

Original post →

More from Research

Research channel →