PufferLib 5.0 RL Results: 19.5k NetHack Score with Random Classes, Multitask Drone Policies
j_foerst · x · 2026-09-14
jsuarez published a rundown of the five coolest RL results in PufferLib 5.0, with @rockt highlighting a 19.5k score in the NetHack Learning Environment using random classes, regularly reaching The Castle.
Highlights shared in the thread:
- Multitask drone policies: a single policy that can both race and arrange into several different formations.
- Robocode champions via self-play: the 2010 tank-programming game and its monster bots were ported to a high-performance C PufferEnv; PufferLib 5.0 can crush any bot it trains against.
The full list spans five results, showcasing PufferLib 5.0's RL training capability in high skill-ceiling environments.
Related event: PufferLib 5.0 Sets New RL Records on Craftax and NetHack(4 posts)→
More from Research
- FlyWire publishes DIY guide to simulate a fly brain with ~160k neurons — patrickmineault · 2026-09-15
- Stanford and MIT paper: the code harness around an LLM can swing benchmark results up to 6x — burkov · 2026-09-15
- Close to a huge math breakthrough, then scooped by AI: what it means for open science — ScottNover · 2026-09-15
- ECCV 2026 paper ART fixes complex makeup transfer, releases first 2K dataset — jiqizhixin · 2026-09-15
- Why mathematicians resist AI proofs — and why it's not just gatekeeping — rbhar90 · 2026-09-15
- CCN2026 GAC debate recording: is the NeuroAI approach inevitable for understanding the brain? — aran_nayebi · 2026-09-15