NVIDIA's LoGRA cuts RL training memory by up to 45.7%, trains 27B model where Adam OOMs
mark_k · x · 2026-10-08
New research from NVIDIA introduces LoGRA, which cuts average RL training memory by up to 45.7% while preserving performance on tested reasoning tasks. The trick: compress learning signals into compact "gradient sketches" and control each update's size to keep training stable. The model still learns via weight updates, just with far less information resident in GPU memory.
They trained a 27B model for over 1,100 steps on a single node with eight H100 GPUs — where the standard Adam setup ran out of memory. The takeaway: fitting more training onto existing hardware means more room to experiment, and efficiency gains like this deserve more attention.
Related event: NVIDIA Unveils LoGRA, Cutting RL Training Memory by Up to 45.7%(2 posts)→
More from Infra
- Surface Laptop Ultra launches Oct 16 starting at $2,599 with Nvidia RTX Spark — tomwarren · 2026-10-08
- Microsoft Surface AI devices ship Friday: RTX Spark Surface Ultra from $2,599, Dev Box $5,999 — ryanshrout · 2026-10-08
- Microsoft's Surface RTX Spark Dev Box opens at $5,999 with 128GB unified memory — The Verge AI · 2026-10-08
- XPU Grasshopper claims AI-co-designed chip, 816x faster in 13 weeks — ycombinator · 2026-10-08
- NSA spending billions of dollars a year testing frontier AI models, sources say — coherence · 2026-10-08
- Hyperscalers will do anything to shave a microcent off pluggable transceiver costs — jwt0625 · 2026-10-08