On two L40S GPUs, vLLM tensor parallelism beat pipeline parallelism in a Qwen3.6-35B FP8 benchmark

amiitk · reddit · 2026-07-29

A Reddit user benchmarked vLLM on a Proxmox VM with two Nvidia L40S cards passed through to Ubuntu and found tensor parallelism consistently beat pipeline parallelism and data parallelism for a Qwen3.6-35B FP8 deployment.

What they ran

Results

The post asks whether the benchmark is misleading or whether TP really is the better choice in this setup, and shares the exact vLLM flags used for TP, DP, and PP experiments.

Related event: Benchmark Shows TP Outperforms PP for Qwen on Dual L40S(2 posts)→

Original post →

More from Infra

Infra channel →