DiG-bench: A Text-Based Benchmark for Evaluating Frontier LLM 'Discovery Capabilities'

AndrewLampinen · x · 2026-08-12

Replied to a discussion about the new DiG-bench. The original post introduced DiG-bench as a novel benchmark that tests frontier AI models via text-based 'discovery games', probing discovery capabilities in the natural domain of language models without visual confounders. The reply sought clarification on the human baseline metrics, specifically asking about the pass-at-K value and the demographic of the human participants tested.

Related event: DiG-bench: 70 Text Games Expose Reasoning Gaps in Frontier Models(13 posts)→

Original post →

More from Research

Research channel →