Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization
Spectra-Global · reddit · 2026-10-06
A team fighting OOM errors while fine-tuning 8B/70B models on consumer GPUs tried a non-quantization route: instead of storing full AdamW optimizer states, they transform gradients with an FFT, drop low-impact frequencies, and compress the state. This preserves directional integrity while cutting VRAM by roughly 50% and allowing much larger batch sizes, at the cost of slightly more compute per step. They avoided 8-bit quantization due to observed convergence degradation. They're sharing an internal Colab environment and baseline weights, inviting others to stress-test the math and discuss other frequency-domain or non-quantization VRAM reduction methods.
More from Infra
- Only 30-40k robots installed in the US last year — the case for self-replicating factories — ihorbeaver · 2026-10-06
- Reflection says Beam hit ~19% MFU with 92% goodput during its 4-week RL run — alexpolozov · 2026-10-06
- SGLang lands layer_boundary for cross-model reuse of TP/DP/CP semantics — BanghuaZ · 2026-10-06
- Scam alert: SpotGPUs.com fakes 'transaction errors' to steal crypto deposits — Equivalent_West7788 · 2026-10-06
- How a ChatGPT-like system actually works: request-to-stream architecture explained — jawadhamza · 2026-10-06
- SGLang Summit 2026 lineup: Intel CEO, Lilian Weng, Perplexity CEO headline SF event — BanghuaZ · 2026-10-06