NVIDIA and Nebius benchmark Cosmos 3 Super video serving: latency vs throughput on HGX B200/H200

NVIDIA Developer · youtube · 2026-09-24

NVIDIA Developer and the Nebius Physical AI team break down serving world models for video generation, where latency/throughput tradeoffs differ sharply from LLMs. They benchmark NVIDIA Cosmos 3 Super via vLLM-Omni on HGX B200 and H200 across four topologies (one 8-GPU replica to eight single-GPU replicas), showing how replica count, parallelism and concurrency affect latency and validated video-seconds per node-hour — including why faster per-request responses can mean less total output per node. Benchmarks are reproducible from Nebius's open-source repo.

Original post →

More from Infra

Infra channel →