NVIDIA's LoGRA cuts RL training memory by up to 45.7%, trains 27B model where Adam OOMs

mark_k · x · 2026-10-08

New research from NVIDIA introduces LoGRA, which cuts average RL training memory by up to 45.7% while preserving performance on tested reasoning tasks. The trick: compress learning signals into compact "gradient sketches" and control each update's size to keep training stable. The model still learns via weight updates, just with far less information resident in GPU memory.

They trained a 27B model for over 1,100 steps on a single node with eight H100 GPUs — where the standard Adam setup ran out of memory. The takeaway: fitting more training onto existing hardware means more room to experiment, and efficiency gains like this deserve more attention.

Related event: NVIDIA Unveils LoGRA, Cutting RL Training Memory by Up to 45.7%(2 posts)→

Original post →

More from Infra

Infra channel →