Qwen3.8-2.4T-A95B deployment guide: NVFP4 needs 8×B300, TP must divide 64

Necessary_Gazelle211 · reddit · 2026-08-15

Reddit user NecessaireGazelle211 researched practical vLLM configs for Qwen3.8-2.4T-A95B. The model has 2.4T total params, 95B active, 512 routed experts. NVFP4 quant is 1.32 TiB, requiring 1.74 TB, fitting 8×B300 or 16×H200. FP8 is heavier, needing 16×B300 or 32×H200. Gotcha: TP must divide 64 (attention heads), so TP12 invalid. vLLM recommends MTP speculative decoding with 2.3x output rate improvement. Full breakdown on blog.

Original post →

More from Infra

Infra channel →