New DiG-bench benchmark: frontier models still stumped by simple discovery tasks

misovalko · x · 2026-08-14

A new benchmark, DiG-bench, from Princeton, MIT, KAUST, and others, tests AI models' ability to discover unknown rules through experimentation in text-based games. Results show frontier models have improved but still fail at surprisingly simple problems in their native text domain.

Original post →

More from Research

Research channel →