Mismatched Tesla V100 pair (16GB+32GB) hits 1,377 prompt tok/s on Qwen3.8 27B locally

OkBase5453 · reddit · 2026-09-20

A Reddit user built a local inference lab from a mismatched pair of Tesla V100-PCIEs (16GB + 32GB, 48GB total) under Proxmox/LXC, benchmarking Qwen3.8 27B Q6KM at 1,376.9 prompt tok/s (2k context), 39.9 decode tok/s, and 1,221.5 pp tok/s at 16k.

Key tuning findings:

The daily-driver recipe: mainline llama.cpp + Qwen3.8 27B Q6 + tensor split + Flash Attention + Q8 KV + NUMA distribute + large batches — a case that old V100s remain surprisingly productive in today's GPU market.

Original post →

More from Infra

Infra channel →