Qwen3.8-27B on 2× RTX 5070 Ti: One vLLM Flag Kills Speed 11×, Full Benchmarks Inside

puthre · reddit · 2026-09-05

The author benchmarks Qwen3.8-27B across llama.cpp, vLLM and NInfer on dual RTX 5070 Ti 16GB (Blackwell, no P2P), with full server configs.

Original post →

More from Infra

Infra channel →