PufferLib 5.0 Doubles SOTA RL Performance on Craftax, Reaching 71% Max Return

jsuarez · x · 2026-09-14

A writeup by @daphnesolves details reaching level 6 in Craftax with tabula rasa on-policy RL: the policy scores an episode return of 162, about 71% of the maximum achievable, learning to craft armor, cast fireball, and defeat numerous enemies.

jsuarez points out that with PufferLib 5.0 the result doubles SOTA performance for RL on Craftax, underscoring the library's continued training-throughput and sample-efficiency gains on sparse-reward, long-horizon environments.

Related event: PufferLib 5.0 Sets New RL Records on Craftax and NetHack(4 posts)→

Original post →

More from Research

Research channel →