Introducing DiG-bench: 70 Text Games to Evaluate AI Scientific Discovery
jcrwhittington · x · 2026-08-12
A research team officially introduced DiG-bench (Discovery Games Benchmark), a benchmark designed to test AI's capacity for scientific discovery.
Core Design:
- Features 70 new interactive text-based games requiring agents to discover and apply unknown rules via experimentation.
- The purely text-based environment eliminates visual confounds, directly testing LLMs' discovery capabilities in their native domain.
- Humans and frontier models are tested under identical interfaces, action sets, and step budgets.
Key Findings:
- Games are categorized into 7 difficulty tiers. All are beatable by humans on their first attempt, but the best AI models struggle significantly with the highest tiers.
- Closed-source models (e.g., Opus 5) significantly outperform open-source models (e.g., Kimi K3).
Related event: DiG-bench: Frontier LLMs Still Stumble on Simple Text Discovery Games(12 posts)→
More from Research
- Another EpochAI Open Problem Bites the Dust — rickasaurus · 2026-08-12
- Harvard Paper: Assigned Roles Alter How Clinical AI Agents Allocate Resources — zakkohane · 2026-08-12
- Are LLM CoTs Unreliable? Researchers Call for Deep Dive into Latent Space Computations — ricklamers · 2026-08-12
- Breaking the Post-Deployment Stagnation: 20+ Startups Bet on Continual Learning — bigdata · 2026-08-12
- Why Self-Distillation Beats GRPO and RLHF in Scaling Continual Learning — AI Engineer · 2026-08-12
- Introducing ContextBench: A LeetCode-Style Playground for Context Engineering — Final_Act_9658 · 2026-08-12