LTX-2.5 on a rented 96GB Blackwell: 10s of 1080p video with audio in 231s for 7 cents
Realistic-Fennel-190 · reddit · 2026-09-23
The author rented an RTX PRO 6000 Blackwell (96GB) for an evening and benchmarked the stock LTX-2.5 text-to-video template in ComfyUI with two runs, no tuning.
Speed and cost
- 5s @ 1280x736: 69.4s total
- 10s @ 1920x1088: 45s pass 1 + 85s pass 2 = 231.5s, about 7 cents of rental time
VRAM insight
- Weights idle at 39GB (21GB transformer + 15GB Gemma text encoder); 10s of 1080p added only 13GB, peaking at 52.5GB
- Takeaway: weights, not clip length/resolution, are what eat your card — a 32GB GPU struggles before generation even starts
Other findings
- Audio and video come from the same latent (ConcatAVLatent/SeparateAVLatent), not a bolted-on soundtrack; requested rain and train horn both landed
- At 1080p, sampling was only 130s of 231s — VAE decode and muxing take nearly half; the refine pass after 2x latent upscale scales worst (5.9x)
- Gotcha: the template's CLIPLoader points to a 4.9GB Gemma file only needed with promptenhance, but ComfyUI statically validates it — download it or Ctrl+B the node
Faces/hands up close and >10s generations remain untested.
More from Infra
- crabbox now runs on boxd: isolated KVM microVMs with ms boot times for repo commands — steipete · 2026-09-23
- AI data centers now demand up to 1,000 MW, a 200x jump in power — Olivier__OG · 2026-09-23
- Rust weight-streaming engine runs FLUX.2 9B on a 12GB RTX 3060, ~9% faster with 2GB resident pool — madtune22 · 2026-09-23
- AMD MI355X Hits 1.7x Perf-Per-Dollar vs DGX B300 via SGLang KV Cache Fix — zephyr_z9 · 2026-09-23
- Apsara Conference's real story: China's full AI stack, chips to robotics — manishkhosiya · 2026-09-23
- Local Qwen3.8-27B Runs Typed Decisions in <10GB VRAM at 170ms — kyr0x0 · 2026-09-23