Running Qwen3.8 at 170K Context on a Single 96GB GPU
UltrMgns · reddit · 2026-08-30
A user shared a practical deployment of Qwen3.8-Flash-Next on a single 96GB GPU, achieving a 170K context window and 110 tokens/sec.
Configuration & Optimizations:
- Used an INT4 quantized version (32GB) via memory mapping.
- Deployed with vLLM, enabling Speculative Decoding (MTP), Prefix Caching, and Chunked Prefill.
- Modified the model.safetensors.index. to drop unnecessary ple-bf16 weights.
- Set --max-model-len 173400, utilizing 89GB VRAM on a single Pro 6000.
Performance:
- The model successfully coded and played complex HTML games in a headless Ubuntu environment, continuously improving its strategy.
- Achieved 87% MTP acceptance rate on code/JSON and >90% prefix cache hit rate on long tasks.
Related event: Qwen3.8 Runs 170K Context on Single 96GB GPU(2 posts)→
More from Infra
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01