Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization

Spectra-Global · reddit · 2026-10-06

A team fighting OOM errors while fine-tuning 8B/70B models on consumer GPUs tried a non-quantization route: instead of storing full AdamW optimizer states, they transform gradients with an FFT, drop low-impact frequencies, and compress the state. This preserves directional integrity while cutting VRAM by roughly 50% and allowing much larger batch sizes, at the cost of slightly more compute per step. They avoided 8-bit quantization due to observed convergence degradation. They're sharing an internal Colab environment and baseline weights, inviting others to stress-test the math and discuss other frequency-domain or non-quantization VRAM reduction methods.

Original post →

More from Infra

Infra channel →