Serving Qwen3.5-9B on 8x T4s: 44 slots at 25-40 tps, where's the bottleneck?
exaknight21 · reddit · 2026-10-10
The author has an Asus ESC4000 G3 (128GB DDR4) running 8x 72W passive NVIDIA T4s, which lack recent fp8 emulation support. Current setup: 8 separate llama.cpp instances serving Qwen3.5-9B-Q5KXL with 44 slots (87k context each) for 10-15 concurrent users, powering a non-coding agentic harness that fetches emails and summarizes documents. Per-slot generation is 25-40 tps with 1400 tps prefill, and latency is a concern. VLLM with AWQ-INT8 on an Mi50 32GB was much faster, so they're asking whether a vLLM fork or quantized build could help — with no budget for new hardware for the next 6-12 months.
More from Infra
- Bittensor GPU rental network sees 45% spend growth, 106% more rentals in monthly report — markjeffrey · 2026-10-10
- Building a pit crew for Grok Bot: frontier model plans, free models grind — alexcovo_eth · 2026-10-10
- AI boom turns into a debt boom: Oracle 5y CDS near record 261bps, implying 20.4% default odds — cyb3rops · 2026-10-10
- Fireworks AI Discloses Security Incident Involving Unauthorized Use of Internal Credentials — lqiao · 2026-10-10
- US grid adds 86GW this year while AI labs need hundreds of GW of power — FinanceYF5 · 2026-10-10
- Running Qwen3.6 35B-A3B with 131K context and vision on a 6GB RTX 2060 — full config — Szadbaverem69 · 2026-10-10