Qwen3.8-2.4T-A95B deployment guide: NVFP4 needs 8×B300, TP must divide 64
Necessary_Gazelle211 · reddit · 2026-08-15
Reddit user NecessaireGazelle211 researched practical vLLM configs for Qwen3.8-2.4T-A95B. The model has 2.4T total params, 95B active, 512 routed experts. NVFP4 quant is 1.32 TiB, requiring 1.74 TB, fitting 8×B300 or 16×H200. FP8 is heavier, needing 16×B300 or 32×H200. Gotcha: TP must divide 64 (attention heads), so TP12 invalid. vLLM recommends MTP speculative decoding with 2.3x output rate improvement. Full breakdown on blog.
More from Infra
- Nvidia Rubin data centers may see 78% net margins, Morgan Stanley reports — Beth_Kindig · 2026-08-15
- Infobip launches Canada data residency in a two-track compliance market — shashib · 2026-08-15
- NVMe drive prices double as hardware infrastructure costs surge — docmilanfar · 2026-08-15
- Paper: Speculative Decoding Often Slower on Mac, Best Case 1.61x — juanviera23 · 2026-08-15
- SpaceX partners with Nvidia on orbital data centers, first satellite launching next year — JOBhakdi · 2026-08-15
- Qdrant + Minima Boost Agentic RAG 2.92x on Single RTX PRO 6000 — qdrant_engine · 2026-08-15