DiG-bench: Frontier Models Still Far Behind Humans at Discovering Hidden Rules

Import AI (Jack Clark) · rss · 2026-08-17

Import AI #469 by Jack Clark covers three items:

1. DiG-bench (Discovery in Games): A new benchmark of 70 handcrafted miniature game worlds where rules and objectives are hidden and must be uncovered through interaction—a text-native cousin of ARC. Most games are kept private to prevent training contamination. Built by researchers from Oxford, Princeton, MIT, KAUST, and others, including Juergen Schmidhuber.

Results across seven difficulty tiers: Opus 5 and Fable 5 with Claude Code lead, followed by GPT-5.5. Only Opus 5 and Fable 5 solved any Tier 7 tasks (0.2); GLM-5.2 and Gemini 3.1 Pro only cleared some Tier 4 levels—while individual humans beat 100% of games. Clark predicts human parity around mid-2027, when recursive self-improvement could seriously kick off.

2. RSI Simulator: Paradigm Research released a Cookie-Clicker-style browser game simulating a company pursuing recursive self-improvement—balancing researchers vs compute, data licensing, and more.

3. Inherent's AI scientist Faraday: A 27B model post-trained on Qwen-3.6-27B supervises a Codex coding agent. Its Replica dataset turns 100 papers (1990-2026) into 310 replication tasks with knocked-out results; Opus 4.7 generates rubrics and a Codex-based judge drives GRPO training. Faraday+Codex beats Opus 4.8 and GPT-5.5 on 73% of in-distribution ML tasks and 60% of held-out AI-for-science tasks—early signs of AI "research taste."

Original post →

More from AGI Musings

AGI Musings channel →