vLLM offloads video decoding to NVDEC, 2x+ throughput on 8×H100
vllm_project · x · 2026-09-22
vLLM integrates PyNvVideoCodec to fix the CPU bottleneck in large-scale video captioning
- Previously vLLM could only decode video via CPU-side OpenCV+FFMPEG; with short outputs (100-200 tokens), decoding dominated and CPU cores maxed out with just 2-4 GPUs
- By offloading decoding to NVIDIA's hardware NVDEC units via PyNvVideoCodec, throughput rises 2x+ on 8×H100 and the CPU bottleneck is gone, scaling well up to 8 GPUs
- Target use cases: AV training video description and searchable metadata; ships with CUDA vLLM releases
More from Infra
- fal's H3 Max generates 5 seconds of frontier-quality video in just 3 seconds — gorkem · 2026-09-22
- NVIDIA's EPD Disaggregation Cuts Multimodal TTFT Up to 5x, E2E Latency 7x — dl_weekly · 2026-09-22
- Running MiniMax H3 locally on a 16GB Mac: 8-10s clips in 15-20 minutes — coberholzer · 2026-09-22
- A Wild Async RL Config: 30 Steps x 25K Rollouts Per Step at Parallelism 4 — willcb · 2026-09-22
- AMD's market cap went from $2B to $1T in the 11 years since Lisa Su became CEO — xiaosun86 · 2026-09-22
- Blackstone poured ~$100B into data centers this year, telling investors this is not the dot-com bubble — JOBhakdi · 2026-09-22