New DiG-bench Tests AI Discovery: Frontier Models Still Stumped by Simple Problems

AccBalanced · x · 2026-08-13

Princeton, MIT, and others release DiG-bench, a text-based benchmark for discovery capabilities. Testing shows frontier models have improved but still fail on surprisingly simple problems.

Related event: DiG-bench Released: 70 Text Games Expose LLMs' Flaws in Scientific Discovery(14 posts)→

Original post →

More from Models

Models channel →