Frontier AI Models Struggle with Simple Discovery: DiG-bench Tests 70 Games
jcrwhittington · x · 2026-08-12
Princeton, MIT, and other institutions introduce DiG-bench, a novel benchmark of 70 text-based interactive games designed to evaluate the scientific discovery capabilities of AI agents.
- Test Design: Each game requires the agent to discover unknown rules through experimentation. Being purely text-based eliminates visual confounds, probing discovery capabilities directly in the natural domain of language models.
- Human vs AI: All games are validated to be beatable by humans on their first attempt. However, despite recent improvements, frontier models still struggle with surprisingly simple problems in the highest difficulty tiers.
- Leaderboard: The platform compares win rates across major models and agentic frameworks (e.g., Opus 5, GPT-5.5, Gemini 3.1 Pro).
Related event: DiG-bench: Frontier LLMs Still Stumble on Simple Text Discovery Games(12 posts)→
More from coding & agent
- AI-Assisted Electronics Design: From SPICE Simulation to EMC — debreuil · 2026-08-12
- Why AI Workflows Still Need FDEs: Non-Deterministic Systems Reshape Deployment — ivory_tang · 2026-08-12
- Garry Tan Open-Sources gstack: Turning Claude Code Into a Full AI Engineering Org — garrytan · 2026-08-12
- Grok 4.6 available in Cursor and API with 2x usage for the first week — Daniel_Farinax · 2026-08-12
- Ref Raises $4M to Fix 'Velocity Sickness' in AI Coding Agents — round · 2026-08-12
- Practical Tip: Run LLM Agent Experiments Free via Kaggle CLI with T4 GPUs — mariofilhoml · 2026-08-12