Counterintuitive VRAM trick: keeping Qwen 3.8 loaded doubles video-gen speed on RTX 5090
Modern_Art_Official · reddit · 2026-09-19
A Reddit user with an RTX 5090 and 96GB RAM running a 5-second 0.98MP video workflow (Dasiwa hybrid model, nvfp4 text encoder, Kijai's experimental VAE, 5 image references) noticed runtimes jumping from 174s to 350-430s when Qwen 3.8 was ejected.
- The culprit is ComfyUI's dynamic vram: even when everything fits in VRAM, the experimental VAE still triggers slow Shared GPU Memory usage (0.7GB spillover when running alone).
- With Qwen left resident, dynamic vram only uses 29.4/76.7GB, keeping runtimes under 3 minutes — at the cost of GPU temps hitting 80°C.
Takeaway: keeping a large model loaded can "occupy" the dynamic vram mechanism and prevent slow paging — a useful counterintuitive config for local video generation.
More from Infra
- Dual RTX 3060 Local LLM Setup: 100k Context at 600 tok/s, $2k Upgrade Paths — gnoremepls · 2026-09-19
- GitHub Next open-sources LocalJev, a local Jev-compatible API built on oMLX and DiffusionGemma — gaganghotra_ · 2026-09-19
- Why does nobody benchmark prefill? Local rigs ignore input processing speed vs API providers — BobtheGodGamer · 2026-09-19
- VanEck: NVDA's bigger risk is customers can't get power; powered-land base case implies ~83% upside — menhguin · 2026-09-19
- Reverse-engineering Claude's subscription limits from unrounded floats: Max 5× is the real sweet spot — RexDouglass · 2026-09-19
- Hyperscaler ROIIC Peaked Near 40% vs 8% Cost of Capital, AI Capex Math Shows — menhguin · 2026-09-19