34GB Model Stack on an 8GB RTX 4060: LTX2.5 Generates 10s Video Locally
BigBullshitta · reddit · 2026-09-17
A Reddit user documents how far local generative AI has come in 18 months, running on just an RTX 4060 8GB + 32GB RAM:
- Previously, one image took 1 min (ZImageTurbo) or 15s (SDXL); video was out of reach.
- Using Wan2GP + ComfyUI with an int8-convrot Krea2 (Klein 2 9B) quant (13GB), 1.2MP images now take 20-30 seconds, with no more VRAM-swap stalls.
- LTX2.5 i2v produced a 10s, 24fps, 768x1024 video in about the time it took to microwave food — and it looked great.
- An accidental 1280x1280 run kept the GPU at 98-100% utilization for 40 minutes, producing the equivalent of 240 1.6MP images from a 34GB model stack.
The author credits recent tensor core / convrot / sage quantization and memory-management advances for letting consumer GPUs handle models far exceeding VRAM.
More from Infra
- DeepSeek-V4.1 Flash deep dive: pushing KV cache compression to the limit at 420 tok/s — teortaxesTex · 2026-09-17
- Four Scheduling Techniques Flatten MoE Training Memory Peaks, Enabling 1M Context at 10.4x Throughput — Shrey Pandit · 2026-09-17
- GLM agent built its own inference infra in two weeks, tripling end-to-end throughput — jietang · 2026-09-17
- Open-source PolyServe autotuner boosts LLM throughput 56–97% across three GPUs — AffectionateSir8341 · 2026-09-17
- China Telecom open-sources Xing4.0-29B MoE, first in class trained fully on Ascend NPUs — Skyline34rGt · 2026-09-17
- Leaked PR Hints Circuit & Chisel's GPU Router Adds "union-alpha" Per-Account Name for Pareto — cephaloform · 2026-09-17