RL agent beats hand-coded AutoAscend on NetHack with PufferLib 5.0, 5x prior SOTA

jsuarez · x · 2026-09-14

jsuarez highlights his summer intern's work: an RL agent trained with PufferLib 5.0 scored >5x the previous SOTA on NetHack, beating the best hand-coded scripted agent (AutoAscend). The linked article explains how RL tackles NetHack's long horizons, partial observability and run-ending random interactions to top the NetHack Challenge.

Related event: PufferLib 5.0 Sets New RL Records on Craftax and NetHack(4 posts)→

Original post →

More from Research

Research channel →