Counterintuitive VRAM trick: keeping Qwen 3.8 loaded doubles video-gen speed on RTX 5090

Modern_Art_Official · reddit · 2026-09-19

A Reddit user with an RTX 5090 and 96GB RAM running a 5-second 0.98MP video workflow (Dasiwa hybrid model, nvfp4 text encoder, Kijai's experimental VAE, 5 image references) noticed runtimes jumping from 174s to 350-430s when Qwen 3.8 was ejected.

Takeaway: keeping a large model loaded can "occupy" the dynamic vram mechanism and prevent slow paging — a useful counterintuitive config for local video generation.

Original post →

More from Infra

Infra channel →