Deep dive into CPU Optimizer Offload: Train 131k context on 32 GPUs

samsja19 · x · 2026-08-26

This feature fully offloads gradients, optimizer states, and steps to the CPU, hiding communication and computation during backward passes. For long sequences where attention dominates, it achieves nearly 100% overlap. While large-scale runs use FSDR, this allows debugging massive models (like GLM5 at 131k context) on just 32 GPUs. The author predicts CPU offloading will become standard during training amid VRAM shortages.

Original post →

More from Infra

Infra channel →