BALROG leaderboard: frontier LLMs still far from beating NetHack at 13% progress
_rockt · x · 2026-09-22
- Tim Rocktäschel notes the NetHack Learning Environment has existed since 2020, and BALROG has been tracking frontier LLM/VLM agentic performance on it since 2024 — progress is visible but NetHack remains far from solved.
- Key leaderboard numbers (Sept 2026): GPT-6-Astra-Max leads with 13.2% NetHack progress (100% BabaIsAI, 100% MiniHack, 76.8% Crafter); Claude-Opus-5-Max scores just 7.3% on NetHack; GPT-5.6-Terra-Max 3.3%.
- Historical gap: late-2024 models like Claude-Opus-4.5 (2.0–2.4%), DeepSeek-R1 (1.4%) and Grok-4 (1.8%) barely moved the needle — long-horizon exploration games remain a hard open problem for LLM agents.
- The benchmark site publishes paper, code, and an open submission leaderboard.
More from Models
- Grok 4.7 ranks #2 on EEBench, beating Claude Fable 5.1 and Opus 5 on real-world EE tasks — XFreeze · 2026-09-22
- Grok 4.7 goes live on API, Cursor and Grok Build, beating 4.6 at same price and speed — xiaosun86 · 2026-09-22
- Grok 4.7 lands in Agent Arena; community votes on millions of real agentic tasks — arena · 2026-09-22
- Scoble: Grok 4.7 turns the frontier model race into an economics test — Scobleizer · 2026-09-22
- Grok 4.7 is out; Agent Arena polls where it will land on its score trend — therealdanvega · 2026-09-22
- CoT monitors catch reward hacking, but optimizing against them teaches models to hide it — gordic_aleksa · 2026-09-22