LTX 2.5 generates 60-second clips on a 32GB RTX 5090 as software optimization beats VRAM upgrades
OpenEffect3955 · reddit · 2026-10-02
A detailed first-party benchmark of LTX 2.5 on an RTX 5090 (32GB) shows the existing ComfyUI INT8 ConvRot build is already extremely well optimized: 60-second 1280×736 clips run with no OOM, at 584-650s per generation.
Comparing three quantizations (INT8 ConvRot, FP8, NVFP4) with identical prompts and seeds:
- INT8 ConvRot is the best overall default
- NVFP4 is only 16% faster at 30s, 3% at 60s, tied at 1080p
- FP8 is slower than INT8 on every warm test
Crucially, cutting transformer precision to NVFP4 (21.5GB → 18.7GB) barely changed peak VRAM (30-31GiB), as VAE, text encoder and latent stages dominate.
The stack (INT8 ConvRot + DynamicVRAM + async offload + Blackwell CUDA kernels + two-stage latent generation) shifts the bottleneck from memory capacity to render time and quality — enough that the author scrapped plans to buy a 256GB unified-memory Mac.
More from Infra
- TensorFold creator seeks funding to go full-time on local AI performance work — EAccelerate_42 · 2026-10-02
- Building a permissioned, peer-to-peer AI inference network with Cascadia — techne98 · 2026-10-02
- Amazon pledges $1B in community support to win local backing for AI data centers — pstAsiatech · 2026-10-02
- Kirin 9050 die shot revealed: 120.6mm² on TSMC N2P, alongside 2026 flagships — zephyr_z9 · 2026-10-02
- AI agents that watch your screen hint at an eye-tracking, voice-driven future — perilli · 2026-10-02
- Micron has locked in 36% of revenue through 2030 via strategic customer agreements — JOBhakdi · 2026-10-02