PufferLib 5.0 hits 60M steps/sec single-GPU RL training, solves Breakout in under a second
jsuarez · x · 2026-09-17
PufferLib 5.0 ships alongside an in-depth engineering writeup: up to 60M useful training steps per second on a single GPU, enough to solve Breakout in under a second. Version 4.0's custom CUDA C stack took an 8-month slog; 5.0 polishes the jank off.
Key numbers: throughput scales linearly down to below 100k parameters — 50M sps at 100k params, 10M sps at 1M params, and 1.8M sps at 10M params. The author argues this reverse-direction scaling (small-model efficiency, not the communication-overhead-besieged scaling-up race) is what makes the result striking.
More from Infra
- VC-Attention: training-free low-bit attention hits 1.9x on B200, beating FlashAttention-4 — xiuyu_l · 2026-09-17
- mlx.fast fixes speed-display bug: MLX kernels hit 80.6 tps, nearing 100% speedup milestone — HankYeomans · 2026-09-17
- Memory shortage hits checkout: Xiaomi raises phone prices 200-1,000 yuan as DRAM stock dips under 10 days — tengyanAI · 2026-09-17
- Burkov Slams OpenRouter Reliability: Fallback Models Fail Together — burkov · 2026-09-17
- Common Crawl puts crawl archives on Hugging Face Storage Bucket, with a getting-started guide — vanstriendaniel · 2026-09-17
- Anthropic Signs A$32B Queensland Data Center Deal, Claude's First Australian Footprint — ocean_protocol · 2026-09-17