FourTune: True 4-bit Diffusion Post-Training
amaarora · x · 2026-07-19
FourTune addresses the fact that **diffusion model post-training doesn't truly become 4-bit just by quantizing weights to 4-bit**, as activations and backward gradients still dominate compute and memory overhead. It proposes three design choices to stabilize and accelerate training: - A frozen stabilizer branch to handle quantization-sensitive outliers; - Block quantization to enable transposed matrix multiplication in 4-bit backpropagation; - Fused kernels to reduce memory traffic from small branches. Experimentally, the authors report that on FLUX customization tasks, the method achieves **2.25x lower memory** and **2.27x faster stepping** compared to BF16 LoRA. The accompanying charts compare training latency, memory footprint, and customization quality across different settings.
More from Multimodal
- Reddit users say Krea 2 Turbo regains strong facial expressions with bypass LoRAs — YentaMagenta · 2026-07-21
- AI music demo blends Suno v5.5, Reason Studios and Grok Imagine 1.5 — Kyrannio · 2026-07-21
- AI anime workflow article breaks storytelling into repeatable prompt steps — Aiden_Tech_Ai · 2026-07-21
- Claude is being pitched as a free workflow for viral YouTube Shorts scripts — Aiden_Tech_Ai · 2026-07-21
- A Bittensor game demo claims a two-person team built a playable 3D world in 30 days — markjeffrey · 2026-07-21
- Gemini Omni Flash turns a boat cabin into a cave inside Flow — chrisfirst · 2026-07-21