DiG-bench: Frontier AI Models Win Only 20% of Simple Discovery Games
Princeton, MIT, KAUST and others released DiG-bench, a benchmark testing scientific discovery via text games where models must uncover unknown rules through experimentation. Even top frontier models win only about 20% of the simplest games.
2026-08-14 ~ 2026-08-15 · 2 related posts
- New DiG-bench benchmark: frontier models still stumped by simple discovery tasks — misovalko · 2026-08-14
1 near-duplicate retellings: misovalko