PufferLib 5.0 Hits 60M Steps/Sec on a Single RTX 5090: How 5 Degrees of Parallelism Work
jsuarez · x · 2026-09-14
jsuarez explains how PufferLib 5.0 broke 60M steps/second of RL training on a single RTX 5090 — months of simulated data per second for most environments.
- Environment counts aren't benchmark inflation: hundreds to thousands of envs per GPU, chosen via a 1000+ experiment hyperparameter sweep balancing task performance vs wall-clock time, with results public in the Constellation dashboard.
- Two major parallelization options: thread-based CPU parallelism and more (post truncated), organized around five degrees of parallelism.
He notes they could exceed 100M SPS with a million environments but report maximum throughput under useful configurations.
More from Infra
- Sandbox tip: bake dependencies into the image instead of pip-installing at runtime — xeophon · 2026-09-14
- US's No.2 law firm Latham & Watkins builds in-house AI stack with Nvidia servers — ayushtweetshere · 2026-09-14
- Oracle Cuts Double-Digit % of Some Teams While Hiring Aggressively for Data Centers and AI — mkheck · 2026-09-14
- Running Two Models Across Strix Halo + r9700 Hits OOM: Full Config Shared — El_90 · 2026-09-14
- Musk: AI will be 99% of SpaceX's value within four to five years — XFreeze · 2026-09-14
- Weaker enterprise HBM demand could finally normalize DRAM and NAND pricing, argues analyst — eyishazyer · 2026-09-14