Benchmarking LingBot-Video: 3s of 1080p in 20 Mins on 4x RTX 6000
NewVeterinarian5384 · reddit · 2026-07-29
A developer tested the video generation model LingBot-Video on a workstation equipped with 4x RTX PRO 6000 Max-Q (96GB VRAM per card).
Hardware & Configuration:
- Setup: 4x RTX PRO 6000 Blackwell Max-Q, 512GB system RAM, PCIe 5, no NVLink.
- Parameters: 1088x1920 resolution, 73 frames, 24 fps (3 seconds duration).
- Pipeline: 480x832 base pass followed by a 1080p refiner, utilizing FSDP2 across 4 cards with context parallelism.
Performance & Details:
- Metrics: End-to-end took just under 20 minutes (refiner took 3/4 of the time), peak VRAM was 57GB per card.
- Engineering: The refiner is a second complete 30B model, doubling the compute at 1080p. The author also noted OOM issues when adapting the default 8-rank script down to 4, which took an evening to resolve using Claude Code.
- Quality: The generated water flow held up well physically, though impact dynamics couldn't be tested as nothing landed in frame.
More from Infra
- OpenRouter’s moat comes from routing data and tooling it can refine multiple times a day — mmurph · 2026-07-29
- Hyperscaler credit spreads may be overpricing risk as GPU spot rents run 2x contract rates — GavinSBaker · 2026-07-29
- OpenAI adds GPT-Live-Transcribe and GPT-Transcribe to its API — OpenAIDevs · 2026-07-29
- NVIDIA shares a guide to training open models with RL on Prime Intellect — NVIDIAAI · 2026-07-29
- Podcast spotlights the future of vector databases in the AI stack — CShorten30 · 2026-07-29
- Qdrant, Future AGI and AWS set a talk on self-improving agent retrieval loops — qdrant_engine · 2026-07-29