Dev streams a 66GB unquantized 22B video model on an iGPU with only 15.6GB shared RAM
Business_Swordfish_5 · reddit · 2026-10-09
An aspiring inference engineer ran LTX-2.5 (22B, bf16, unquantized, with audio) on an Intel Core Ultra 5 iGPU with 15.6GB shared RAM by never loading the 66GB model: weights stay on NVMe and stream layer-by-layer with aligned reads, a pinned buffer, and a prefetch thread. Timings: 14 min for a 5.4s 1536x896 anime clip with sound, 47 min at best-quality 24fps, 20 min for a 4s realistic clip. Known issues include motion smearing and 12fps anime output. Full bug log (memory alignment, bf16 vocoder silence) in the open-source repo.
More from Infra
- GPUs already within 2x of brain efficiency, and still beat human workers on energy — MikePFrank · 2026-10-09
- Do AI agents still need Kubernetes? Berlin event says yes, with agent-on-K8s cases — Al_Grigor · 2026-10-09
- Cloud Backlogs Hit $1.69T, CoreWeave Posts $2.58B Quarter as Inference Becomes the Battleground — FinanceYF5 · 2026-10-09
- NVIDIA Is AI's Central Bank: A100 Paper Citations Still Beat H100+H200 Combined — FinanceYF5 · 2026-10-09
- Browser-Based Calculator Crunches the Real Power Cost of Self-Hosted LLMs vs Cloud APIs — paq85 · 2026-10-09
- PartyKit shuts down free hosted platform 2.5 years after Cloudflare acquisition — threepointone · 2026-10-09