8GB VRAM runs Flux and Wan 2.2 locally fine: three wrong settings, not the GPU, were the bottleneck
leonbuilds · reddit · 2026-09-08
On an 8GB RTX 5060 laptop (WSL2), the author runs Flux and Wan 2.2 fully locally. The bottleneck was settings, not the card: Flux fp4 (svdq-fp4r32 via nunchaku) at 6.6GB fits fully in VRAM at 3.4-3.8s/step vs 11-22s/step for the smaller Q4KS GGUF that offloads 1127MB; cpuoffload=auto dropped GPU usage to 1% at 234s/step; and Wan ti2v-5b q5 runs 960x544, 81 frames at 1.95s/step with 3GB free — far beyond the 480x832 the author assumed was the ceiling.
More from Infra
- Weaviate 1.39's 4-bit Rotational Quantization cuts 10M OpenAI embeddings from 61GB to 7.8GB RAM — philipvollet · 2026-09-08
- Running qwen3-tts-1.7b Locally via llama.cpp + Vulkan Bluescreens a Tablet — NERDDISCO · 2026-09-08
- MCP testing: one question went from 437k to 78k input tokens — cost is driven by round trips — loookashow · 2026-09-08
- Japan plans $60 billion investment in data centers — JosephJacks_ · 2026-09-08
- Vaire Computing CTO named MIT TR35 for chips that recycle waste heat into energy — MikePFrank · 2026-09-08
- GMKtec's EVO-X5 Pro mini PC claims to run a 300B LLM fully offline — Dave_Maynor · 2026-09-08