NVIDIA and Nebius benchmark Cosmos 3 Super video serving: latency vs throughput on HGX B200/H200
NVIDIA Developer · youtube · 2026-09-24
NVIDIA Developer and the Nebius Physical AI team break down serving world models for video generation, where latency/throughput tradeoffs differ sharply from LLMs. They benchmark NVIDIA Cosmos 3 Super via vLLM-Omni on HGX B200 and H200 across four topologies (one 8-GPU replica to eight single-GPU replicas), showing how replica count, parallelism and concurrency affect latency and validated video-seconds per node-hour — including why faster per-request responses can mean less total output per node. Benchmarks are reproducible from Nebius's open-source repo.
More from Infra
- Why the Semiconductor Industry Is Betting on Optical Interconnects for AI — BenBajarin · 2026-09-24
- Qualcomm Officially Brings Linux to Snapdragon X2 Chips, Debian Coming End of This Year — tomwarren · 2026-09-24
- PyTorchCon 2026 Poster to Show Warm-Start Autotuning for Helion GPU Kernel DSL — PyTorch · 2026-09-24
- Qualcomm Announces Linux Support for Snapdragon X2 Elite — karlfreund · 2026-09-24
- MediaTek's consensus growth seen at 65-85% on TPU demand, analyst says — BenBajarin · 2026-09-24
- Company From 'Gas to Gigawatts' Expert Interview Lands on UBS Most Preferred List — BenBajarin · 2026-09-24