PufferAI Expert: RL Bottleneck Was 1000x Slow Code, Not Algorithms
ziv_ravid · x · 2026-07-31
A podcast interview featuring Joseph Suarez from PufferAI diving deep into the core bottlenecks of Reinforcement Learning (RL).
Key Insights & Discussions:
- Algorithms Weren't the Problem: Suarez argues RL never had an algorithm problem. Instead, the code was universally about 1000x too slow. By fixing this, problems that previously took months can now be solved in seconds on a single GPU.
- Simulators & Compute: Explores what makes a good simulator for RL and explains why most of PufferAI's sims run on CPUs rather than GPUs.
- Skepticism of World Models: He expresses doubt about the current trend of using world models as environment generators for RL training.
- Future Applications: Outlines the vision to apply this highly efficient RL framework to material science and biological simulations.
More from Infra
- Jeff Dean: AI Models Are Already Junior Engineers, Inference Hardware is Next — ycombinator · 2026-07-31
- Viewpoint: Low Interest Rates Indicate We Are Not Overinvesting in AI Compute — tszzl · 2026-07-31
- Optimizing LLM Inference TPS on NVIDIA Blackwell Is the Funnest Thing — abhijithneil · 2026-07-31
- Cloud Giants Scale Up: Google Cloud Revenue Surges 82% YoY — davidyin44 · 2026-07-31
- Running Krea2 on RTX 5090: ComfyUI Setup Faces VRAM Bottlenecks — orangeflyingmonkey_ · 2026-07-31
- AI Capex Gap: Apple's 9-Month Spend Less Than Amazon's Single Quarter — mkheck · 2026-07-31