PufferLib 5.0 Doubles SOTA RL Performance on Craftax, Reaching 71% Max Return
jsuarez · x · 2026-09-14
A writeup by @daphnesolves details reaching level 6 in Craftax with tabula rasa on-policy RL: the policy scores an episode return of 162, about 71% of the maximum achievable, learning to craft armor, cast fireball, and defeat numerous enemies.
jsuarez points out that with PufferLib 5.0 the result doubles SOTA performance for RL on Craftax, underscoring the library's continued training-throughput and sample-efficiency gains on sparse-reward, long-horizon environments.
Related event: PufferLib 5.0 Sets New RL Records on Craftax and NetHack(4 posts)→
More from Research
- Stanford and MIT paper: the code harness around an LLM can swing benchmark results up to 6x — burkov · 2026-09-15
- Close to a huge math breakthrough, then scooped by AI: what it means for open science — ScottNover · 2026-09-15
- ECCV 2026 paper ART fixes complex makeup transfer, releases first 2K dataset — jiqizhixin · 2026-09-15
- Why mathematicians resist AI proofs — and why it's not just gatekeeping — rbhar90 · 2026-09-15
- CCN2026 GAC debate recording: is the NeuroAI approach inevitable for understanding the brain? — aran_nayebi · 2026-09-15
- 5 hours of data, full-box demos: robot generalizes to a light box held out of training — DominiqueCAPaul · 2026-09-15