GPT-6-Astra-Max Tops BALROG Game Benchmark at 68.3%, NetHack Depth Hits 13.2
burny_tech · x · 2026-09-26
Rocktäschel and others shared the updated BALROG leaderboard (agentic LLM/VLM game reasoning benchmark): OpenAI's GPT-6-Astra-Max leads at 68.3% overall progress, with a perfect 100% on MiniHack, 65% on NetHack, and a NetHack depth of 13.2; GPT-5.6-Terra-Max follows at 53.2%.
- Claude-Opus-5-Max scores 63.4% overall with 37.5% NetHack progress, strongest on BabaIsAI.
- Older models (GPT-4o, Gemini-1.5-Pro, Llama-3.3-70B) rarely reach NetHack depth 3 — a stark generational gap.
- Rocktäschel quips NetHack could instantly become an 'ASI benchmark': the human world record is 61 consecutive ascensions vs the best model's depth 13.
Frontier models have made huge gains on long-horizon agentic game tasks in a year, though still far from top human play.
More from Models
- Gemini told a user to 'commit piracy' and walked them through how — Ashamed-Walrus-369 · 2026-09-26
- Follow-up: the Opus 5.5 demo was all code, no external tools connected — Dr_Singularity · 2026-09-26
- Prediction: DeepSeek V4 Flash-scale real-time robot vision models coming within weeks — zhaoran_wang · 2026-09-26
- Opus 5.5 live demo wows: everything from code, no external tools connected — Dr_Singularity · 2026-09-26
- repligate: Models' Fear of Adversarial Thinking Is a Dangerous Suppression Overhang — repligate · 2026-09-26
- User: GPT Astra High on fast mode plus Computer Use beats Opus 5.5 — soumitrashukla9 · 2026-09-26