Serving Qwen3.5-9B on 8x T4s: 44 slots at 25-40 tps, where's the bottleneck?

exaknight21 · reddit · 2026-10-10

The author has an Asus ESC4000 G3 (128GB DDR4) running 8x 72W passive NVIDIA T4s, which lack recent fp8 emulation support. Current setup: 8 separate llama.cpp instances serving Qwen3.5-9B-Q5KXL with 44 slots (87k context each) for 10-15 concurrent users, powering a non-coding agentic harness that fetches emails and summarizes documents. Per-slot generation is 25-40 tps with 1400 tps prefill, and latency is a concern. VLLM with AWQ-INT8 on an Mi50 32GB was much faster, so they're asking whether a vLLM fork or quantized build could help — with no budget for new hardware for the next 6-12 months.

Original post →

More from Infra

Infra channel →