GPT-6-Astra-Max Tops BALROG Game Benchmark at 68.3%, NetHack Depth Hits 13.2

burny_tech · x · 2026-09-26

Rocktäschel and others shared the updated BALROG leaderboard (agentic LLM/VLM game reasoning benchmark): OpenAI's GPT-6-Astra-Max leads at 68.3% overall progress, with a perfect 100% on MiniHack, 65% on NetHack, and a NetHack depth of 13.2; GPT-5.6-Terra-Max follows at 53.2%.

Frontier models have made huge gains on long-horizon agentic game tasks in a year, though still far from top human play.

Original post →

More from Models

Models channel →