DiG-bench: 70 Text-Based Discovery Games to Test AI's Scientific Exploration
jcrwhittington · x · 2026-08-12
To ensure AI can genuinely make new scientific discoveries rather than just combining known tools, researchers introduced DiG-bench, a benchmark of 70 handcrafted text-based discovery games.
- Gameplay: Each game is a self-contained world with hidden rules and win conditions, requiring real experimentation and exploration to beat.
- Difficulty: Features seven difficulty tiers calibrated to frontier models. The games are challenging but beatable, with at least one human beating every game on their first try.
- Data Split: 21 games are public, while 49 remain private for secure and reliable evaluation.
Related event: DiG-bench: 70 Text Games Expose Reasoning Gaps in Frontier Models(13 posts)→
More from Research
- Eric Jang: Building an Automated Researcher with Claude Code to Rewrite AlphaGo — michaelrzhang · 2026-08-13
- Deep Dive with Yulu Gan on How Pretraining Shapes Weight Distributions — yacinelearning · 2026-08-13
- UCLA Team Publishes Metabolic Atlas of Human Cortex, Revealing Glucose Metabolism Controls Cell Fate — anne_churchland · 2026-08-13
- A Single Scaling Curve Is Not a Scaling Law, Caution in AI Research — Majumdar_Ani · 2026-08-13
- Redwood and Anthropic Launch Conceptual Reasoning Index to Evaluate AI Safety Reasoning — RyanGreenblatt · 2026-08-13
- Ex-OpenAI Researcher: Human Data Labeling Isn't the Main Bottleneck for AI Progress — RyanGreenblatt · 2026-08-13