NetHack benchmark reportedly solved after years of slow LLM progress
PMinervini · x · 2026-09-25
Reports on X suggest the NetHack Learning Environment (NLE) benchmark may have finally been "solved." NLE has been open since 2020, and since 2024 it has been used to evaluate frontier LLMs, with progress long described as slow and far from complete. The "solved" claim comes via a surprised retweet; details remain to be verified.
More from Models
- Redditor: models are now good enough that I stopped caring about prompt engineering — Large-Excitement6573 · 2026-09-25
- User says OpenAI's Sol 6 wastes their time, switches to $100 Claude plan and prefers it — AirportEither2456 · 2026-09-25
- Nagi-ENORMOUS beats Jev, Semif, Laya on Game Arena Benchmark — No_Skill_8393 · 2026-09-25
- Principia: video models know less high-school physics than you'd think — mariyaivasileva · 2026-09-25
- Post-scarcity intelligence defined: Opus 5.5-level smarts at $0.10 input / $0.50 output per million tokens — Sect-Sister-Reads-42 · 2026-09-25
- 'Capability gaslighting' is real—but not in agentic engineering, where code is consistent — PawelHuryn · 2026-09-25