Dev streams a 66GB unquantized 22B video model on an iGPU with only 15.6GB shared RAM

Business_Swordfish_5 · reddit · 2026-10-09

An aspiring inference engineer ran LTX-2.5 (22B, bf16, unquantized, with audio) on an Intel Core Ultra 5 iGPU with 15.6GB shared RAM by never loading the 66GB model: weights stay on NVMe and stream layer-by-layer with aligned reads, a pinned buffer, and a prefetch thread. Timings: 14 min for a 5.4s 1536x896 anime clip with sound, 47 min at best-quality 24fps, 20 min for a 4s realistic clip. Known issues include motion smearing and 12fps anime output. Full bug log (memory alignment, bf16 vocoder silence) in the open-source repo.

Original post →

More from Infra

Infra channel →