Benchmarking Qwen3.6-35B on Dual L40s: Tensor Parallelism Beats Pipeline and Data Parallelism
gulensah · reddit · 2026-07-29
A user benchmarks Qwen3.6-35B-A3B-FP8 on two Nvidia L40s (no NVLink) in a Proxmox VM, comparing Tensor Parallelism (TP), Pipeline Parallelism (PP), and Data Parallelism (DP). Results show TP outperforms PP and DP in requests, latency, and throughput, contradicting vLLM docs that recommend PP without NVLink. Includes config and benchmark data.
Related event: Benchmark Shows TP Outperforms PP for Qwen on Dual L40S(2 posts)→
More from Infra
- Why AI workloads can make renting cloud compute cost more than owning it — DavidLinthicum · 2026-07-29
- A CLI tool now ranks multi-agent runs by cost and flags repeated context tokens — rrk059 · 2026-07-29
- Colibri-based runner brings Kimi K3 GGUFs to workstation-class hardware — Responsible_Fig_1271 · 2026-07-29
- ML teams still pin torch and Hugging Face deps to keep runs reproducible — vboykis · 2026-07-29
- Jensen Huang Reportedly Committed $4 Billion in Compute to Scale Ilya's AI — iruletheworldmo · 2026-07-29
- AI labs may be hitting compute limits as next leap gets more expensive — haider1 · 2026-07-29