DiG-bench: Evaluating AI Discovery via 70 Hidden-Rule Text Games
jcrwhittington · x · 2026-08-12
To evaluate AI's genuine ability to explore and make scientific discoveries in unknown environments, researchers introduced DiG-bench. The benchmark consists of 70 self-contained text-based discovery games, each rendered as a minimal string where the AI acts as an arrow navigating the world.
Key features include:
- Hidden Rules: The rules and win conditions are completely concealed, forcing the AI to rely on real experimentation and exploration to progress.
- Calibrated Difficulty: Features 7 difficulty tiers aligned with frontier models. The games are hard but beatable, with every game having been solved by at least one human on their first try.
- Secure Evaluation: 21 games are public, while 49 remain private to prevent contamination and ensure secure evaluation.
The author argues that for AI to chart new science rather than just combining known tools, it must first prove it can make discoveries in the easiest possible settings.
Related event: DiG-bench: 70 Text Games Expose Reasoning Gaps in Frontier Models(13 posts)→
More from Research
- Eric Jang: Building an Automated Researcher with Claude Code to Rewrite AlphaGo — michaelrzhang · 2026-08-13
- Deep Dive with Yulu Gan on How Pretraining Shapes Weight Distributions — yacinelearning · 2026-08-13
- UCLA Team Publishes Metabolic Atlas of Human Cortex, Revealing Glucose Metabolism Controls Cell Fate — anne_churchland · 2026-08-13
- A Single Scaling Curve Is Not a Scaling Law, Caution in AI Research — Majumdar_Ani · 2026-08-13
- Redwood and Anthropic Launch Conceptual Reasoning Index to Evaluate AI Safety Reasoning — RyanGreenblatt · 2026-08-13
- Ex-OpenAI Researcher: Human Data Labeling Isn't the Main Bottleneck for AI Progress — RyanGreenblatt · 2026-08-13