NVIDIA's LoGRA cuts LLM RL training memory by up to 45.7%, enables 27B RL on a single 8-GPU node
nvidia · hf · 2026-10-07
NVIDIA introduces LoGRA, a memory-efficient approach to LLM RL post-training:
- Core idea: retain useful learning signals in low-rank gradient sketches, supporting both model updates and efficient policy synchronization.
- Stability: paired with predicted-KL step control that estimates policy changes before each update and adjusts step size to prevent destabilizing jumps.
- Results: reduces average training memory by up to 45.7% on reasoning tasks without sacrificing performance; enables stable training of a 27B model for over 1,100 steps on a single 8-GPU node, where dense Adam runs out of memory.
- Code is open-sourced in the Molt library.
More from Infra
- Intel exec: agentic AI spends real time waiting on business systems, end-to-end wall-clock is the metric that matters — ryanshrout · 2026-10-07
- Serving & Monitoring Local LLMs Across Three Mixed-GPU Machines Without Duct Tape — ziyaulhuk12 · 2026-10-07
- Disaggregated inference is the future, says e/acc's Beff Jezos after panel with General Compute — beffjezos · 2026-10-07
- What Does 'Owning Your Own AI' Technically Mean? Locally Running Open Weights Debated — Imaginary_Choice_430 · 2026-10-07
- Ex-NVIDIA engineer tells the story of testing a chip with a broken memory controller — blelbach · 2026-10-07
- Follow the money risk: memory suppliers take deposits, AI chip makers finance their customers — tengyanAI · 2026-10-07