RL agent beats hand-coded AutoAscend on NetHack with PufferLib 5.0, 5x prior SOTA
jsuarez · x · 2026-09-14
jsuarez highlights his summer intern's work: an RL agent trained with PufferLib 5.0 scored >5x the previous SOTA on NetHack, beating the best hand-coded scripted agent (AutoAscend). The linked article explains how RL tackles NetHack's long horizons, partial observability and run-ending random interactions to top the NetHack Challenge.
Related event: PufferLib 5.0 Sets New RL Records on Craftax and NetHack(4 posts)→
More from Research
- Stanford and MIT paper: the code harness around an LLM can swing benchmark results up to 6x — burkov · 2026-09-15
- Close to a huge math breakthrough, then scooped by AI: what it means for open science — ScottNover · 2026-09-15
- ECCV 2026 paper ART fixes complex makeup transfer, releases first 2K dataset — jiqizhixin · 2026-09-15
- Why mathematicians resist AI proofs — and why it's not just gatekeeping — rbhar90 · 2026-09-15
- CCN2026 GAC debate recording: is the NeuroAI approach inevitable for understanding the brain? — aran_nayebi · 2026-09-15
- 5 hours of data, full-box demos: robot generalizes to a light box held out of training — DominiqueCAPaul · 2026-09-15