Running Full Flux Model on 12GB VRAM: A Local Inference Practice

TheRealFutaFutaTrump · reddit · 2026-08-05

A developer shared their experience running the full Flux model locally using a 12GB RTX 3060 and a 24GB Tesla P40.

Using a dual-boot Linux environment and ComfyUI, the author built a workflow that generates a base image, upscales it, and passes it again for detail enhancement at 3440x1440. By optimizing VRAM allocation (switching offloading to the 3060), render times were drastically reduced from an initial seven and a half hours to under 15 minutes. This demonstrates that older, budget-friendly hardware combinations can comfortably handle local inference for large image models.

Original post →

More from Infra

Infra channel →