Benchmarking Qwen3.6-35B on Dual L40s: Tensor Parallelism Beats Pipeline and Data Parallelism

gulensah · reddit · 2026-07-29

A user benchmarks Qwen3.6-35B-A3B-FP8 on two Nvidia L40s (no NVLink) in a Proxmox VM, comparing Tensor Parallelism (TP), Pipeline Parallelism (PP), and Data Parallelism (DP). Results show TP outperforms PP and DP in requests, latency, and throughput, contradicting vLLM docs that recommend PP without NVLink. Includes config and benchmark data.

Related event: Benchmark Shows TP Outperforms PP for Qwen on Dual L40S(2 posts)→

Original post →

More from Infra

Infra channel →