DiG-bench Released: New Text-Based Benchmark Stumps Frontier LLMs on Simple Discovery
misovalko · x · 2026-08-12
Researchers from Princeton, MIT, and other institutions have introduced DiG-bench (Discovery in Games), a new text-based benchmark for in-context discovery.
Benchmark Mechanics
- The benchmark features novel text-based concepts and mechanisms where players and models must figure out the rules through interaction with very little prior instruction.
- The design draws inspiration from early tests of abstraction and analogy by Doug Hofstadter and Melanie Mitchell.
Key Findings
- While frontier models have improved significantly in recent months, they are still stumped by surprisingly simple discovery problems within their native text domain that humans can easily solve.
Related event: DiG-bench: 70 Text Games Expose Reasoning Gaps in Frontier Models(13 posts)→
More from Research
- Deep Dive with Yulu Gan on How Pretraining Shapes Weight Distributions — yacinelearning · 2026-08-13
- UCLA Team Publishes Metabolic Atlas of Human Cortex, Revealing Glucose Metabolism Controls Cell Fate — anne_churchland · 2026-08-13
- A Single Scaling Curve Is Not a Scaling Law, Caution in AI Research — Majumdar_Ani · 2026-08-13
- Redwood and Anthropic Launch Conceptual Reasoning Index to Evaluate AI Safety Reasoning — RyanGreenblatt · 2026-08-13
- Ex-OpenAI Researcher: Human Data Labeling Isn't the Main Bottleneck for AI Progress — RyanGreenblatt · 2026-08-13
- Ilya's SSI Pushes TTT Paradigm for Real-Time Model Learning — iruletheworldmo · 2026-08-13