GPT Astra 6 hits new BALROG heights with 13% NetHack progression, still "not AGI"
_rockt · x · 2026-09-19
BALROG's leaderboard adds three new entries: GPT 5.6 Sol at max effort remains within error of Gemini 3 and 3.1 Pro, while GPT Astra 6 at max reasoning effort reaches new heights, posting 13% average progression on the NetHack Learning Environment. Researcher rockt comments: "Great progress. Not yet AGI though."
More from Models
- Cactus Releases Needle 3: an 8-29MB On-Device Model Built for Tool Calls — airesearch12 · 2026-09-19
- Nearly 20 openjev models cataloged as hobbyist preps first community leaderboard — airesearch12 · 2026-09-19
- Dev Slams New DeepSeek Model as Distilled Claude Without the Intelligence — Aryvyo · 2026-09-19
- Xiaomi's AI Persona Teased: Unified Model Xiaomi MiMo Launches Tomorrow — xiaohu · 2026-09-19
- Steering Vectors as a Softer Way to Limit Reasoning Budgets? — maddie-lovelace · 2026-09-19
- Two Prompts That Expose How ChatGPT Quietly Rewrites Your Claims — KazTheMerc · 2026-09-19