Running Qwen 3.8 Next on six V100s: MTP nearly doubles output to 43 tok/s
Odd_Caterpillar_2994 · reddit · 2026-09-20
A Redditor shares a full writeup of getting Qwen 3.8 Next running on a 6x V100 rig (TP2 PP3):
- Troubleshooting: one card dropped to PCIe Gen1 x16 (8 hours to fix); sglang-v100 kept OOMing and pxa errored, before 1cat-vllm ran stable.
- Config: speculative decoding capped at 1 due to memory; 8.78 GiB KV cache with 531K tokens.
- Benchmarks: prefill peaks at 4,679 tok/s at 16K input; enabling MTP lifts generation from an average 22.59 to 42.70 tok/s — nearly 2x.
- Thermals: 20-minute gpu-burn stress test stabilizes at 64°C with fans at only 76%.
More from Infra
- $18bn of loans tied to Oracle's New Mexico data centre slide into stressed territory — GaryMarcus · 2026-09-20
- Engineer cuts latency 10x with obscure 'parallel minibatching' trick — unironictechbro · 2026-09-20
- Huawei's Ascend 960DT is its first HBM accelerator: 288GB, 9.6TB/s, 4 PFLOPS FP4 — pstAsiatech · 2026-09-20
- Rumor: Anthropic trails OpenAI in training compute intensity and inference economics — teortaxesTex · 2026-09-20
- Ben Bajarin: Agentic AI will spawn an 'agentic native' CPU tier in datacenters — BenBajarin · 2026-09-20
- Qwen 3.8 Next Flash at 3.05bpw EXL3 runs like Q8 on 3x RTX 3090s, dev reports — nicholas_the_furious · 2026-09-20