DiG-bench: A Text-Based Benchmark for Evaluating Frontier LLM 'Discovery Capabilities'
AndrewLampinen · x · 2026-08-12
Replied to a discussion about the new DiG-bench. The original post introduced DiG-bench as a novel benchmark that tests frontier AI models via text-based 'discovery games', probing discovery capabilities in the natural domain of language models without visual confounders. The reply sought clarification on the human baseline metrics, specifically asking about the pass-at-K value and the demographic of the human participants tested.
Related event: DiG-bench: 70 Text Games Expose Reasoning Gaps in Frontier Models(13 posts)→
More from Research
- Deep Dive with Yulu Gan on How Pretraining Shapes Weight Distributions — yacinelearning · 2026-08-13
- UCLA Team Publishes Metabolic Atlas of Human Cortex, Revealing Glucose Metabolism Controls Cell Fate — anne_churchland · 2026-08-13
- A Single Scaling Curve Is Not a Scaling Law, Caution in AI Research — Majumdar_Ani · 2026-08-13
- Redwood and Anthropic Launch Conceptual Reasoning Index to Evaluate AI Safety Reasoning — RyanGreenblatt · 2026-08-13
- Ex-OpenAI Researcher: Human Data Labeling Isn't the Main Bottleneck for AI Progress — RyanGreenblatt · 2026-08-13
- Ilya's SSI Pushes TTT Paradigm for Real-Time Model Learning — iruletheworldmo · 2026-08-13