PufferLib 5.0: Five Coolest RL Results, from Multitask Drone Policies to Selfplay Robocode Champions
jsuarez · x · 2026-09-14
jsuarez rounds up the five most impressive RL results in PufferLib 5.0:
- Multitask drone policies: a single policy can both race and arrange into multiple formations, built by an intern.
- Robocode champions via selfplay: the team ported the classic tank-coding game and its strongest historical bots to a high-performance C PufferEnv. The standout result: agents trained purely via selfplay, never facing the reference bots directly, still beat every bot they train against.
- Craftax SOTA (post truncated).
The writeup highlights PufferLib's selfplay generalization in high skill-ceiling environments and multitask policy training.
More from Research
- Pharma consortium trains OpenFold3 on 20,000 proprietary structures, beats Boltz-2 with 50% accuracy — TheMoonMidas · 2026-09-14
- Cognichip unveils physics-based AI model claiming up to 100X faster chip design — karlfreund · 2026-09-14
- Daniel Lemire shares his take on AI and math in new video — lemire · 2026-09-14
- AI trained on 20,000 pharma-secret protein structures beats public-data AlphaFold models — MoAlQuraishi · 2026-09-14
- Researcher: Astra Robot Shows Step-Level Image-Action Understanding Unlike Prior VLAs — YuXiang_IRVL · 2026-09-14
- Google's open-sourced fruit fly brain went viral, but nobody can explain how to train it — Haghiri75 · 2026-09-14