AMD 7900 XTX user trains FLUX.2 LoRA on 24GB VRAM with native ROCm
elderon_echar · reddit · 2026-07-24
A detailed guide to training a 32B FLUX.2 LoRA on an AMD RX 7900 XTX 24GB card under native ROCm on Windows, with no CUDA shim.
- The author managed to keep the 32B FLUX.2 dev transformer resident on the GPU and train it with only 32GB of system RAM.
- The post walks through a long chain of failures and fixes: manual safetensors loading, CPU-side quantization, releasing the transformer state dict before loading the text encoder, avoiding layer offloading deadlocks, disabling in-training sampling, enabling expandablesegments, and saving a quantized state dict for faster reloads.
- It also explains how to tell whether training is really using the GPU on ROCm/Windows, because cross-process VRAM checks can be misleading.
- For evaluation, the author renders checkpoints in ComfyUI + ComfyUI-GGUF with GGUF models and says resolution is the main speed knob.
More from Infra
- Intel says it has 10+ long-term foundry customers and is boosting capex — BenBajarin · 2026-07-24
- Intel Commits to High-Volume 14A Production Ramp in 2028 — BenBajarin · 2026-07-24
- Pydantic AI says one agent can spawn 40 test suites, and the laptop fans prove it — AAAzzam · 2026-07-24
- Fluidstack raises $830M Series A at a $7.5B valuation — MxMnr · 2026-07-24
- Google’s capex guide now exceeds all spending from founding through 2021 — ivan_bezdomny · 2026-07-24
- Intel says advanced packaging backlog is building, with billions in sight — BenBajarin · 2026-07-24