LTX 2.5 generates 60-second clips on a 32GB RTX 5090 as software optimization beats VRAM upgrades

OpenEffect3955 · reddit · 2026-10-02

A detailed first-party benchmark of LTX 2.5 on an RTX 5090 (32GB) shows the existing ComfyUI INT8 ConvRot build is already extremely well optimized: 60-second 1280×736 clips run with no OOM, at 584-650s per generation.

Comparing three quantizations (INT8 ConvRot, FP8, NVFP4) with identical prompts and seeds:

Crucially, cutting transformer precision to NVFP4 (21.5GB → 18.7GB) barely changed peak VRAM (30-31GiB), as VAE, text encoder and latent stages dominate.

The stack (INT8 ConvRot + DynamicVRAM + async offload + Blackwell CUDA kernels + two-stage latent generation) shifts the bottleneck from memory capacity to render time and quality — enough that the author scrapped plans to buy a 256GB unified-memory Mac.

Original post →

More from Infra

Infra channel →