New DiG-bench Tests AI Discovery: Frontier Models Still Stumped by Simple Problems
AccBalanced · x · 2026-08-13
Princeton, MIT, and others release DiG-bench, a text-based benchmark for discovery capabilities. Testing shows frontier models have improved but still fail on surprisingly simple problems.
More from Models
- Leaked Grok 4.6 Leads in Agent Tasks, Beats GPT-5.6 in Coding — rohanpaul_ai · 2026-08-13
- Rumor: Claude to Embed Invisible Watermarks in All Outputs — AccBalanced · 2026-08-13
- User confused: Why do I have Gemini Pro access with only AI Plus subscription? — SternButFaiir · 2026-08-13
- Ollama Confirms Computer Use Feature is in the Works — reach_vb · 2026-08-13
- DeepSeek V4 Pro Mocked for Scoring Only 1 Point Higher Than Flash — ns123abc · 2026-08-13
- Developer Predicts Grok Will Win Claude Users Citing xAI's Insane Shipping Speed — XFreeze · 2026-08-13