DiG-bench: Top Frontier Models Clear Only 20% on Simple Discovery Games
misovalko · x · 2026-08-15
Researchers from Princeton, MIT, and others released DiG-bench, a new benchmark designed to evaluate AI discovery capabilities. Using text-based discovery games to probe models in their native domain (avoiding confounding visual elements), the study reveals that while frontier models have improved significantly over the last few months, they remain stumped by surprisingly simple problems. Notably, humans can beat these games on the first try, whereas the best model clears only 20%. The research also found that advanced agentic harnesses provided zero lift over a basic setup.
Related event: DiG-bench: Frontier AI Models Win Only 20% of Simple Discovery Games(2 posts)→
More from Research
- Google Open-Sources HEIR Compiler for Homomorphic Encryption, Enabling Inference on Encrypted Data — DynamicWebPaige · 2026-08-15
- AI helps mathematicians disprove 30-year-old conjecture — skdh · 2026-08-15
- Google open-sources homomorphic encryption compiler for secure inference — DynamicWebPaige · 2026-08-15
- Debate on using AI in academic peer review and disclosure norms — lpachter · 2026-08-15
- Claude solves open stochastic thermodynamics problem — lpachter · 2026-08-15
- GPU contest: batched compact-Householder QR kernel achieves 232x speedup — petrusenko_max · 2026-08-15