Qwen3.8-Flash-Next NVFP4 runs full 256K context on 4x V100 in SGLang-V100
Primary_Exchange21 · reddit · 2026-08-30
The community project SGLang-V100 now supports RadixArk/Qwen3.8-Flash-Next-NVFP4, running full context on 4x V100 32GB with >50GB of ngram offloaded to system RAM. Prefill holds around 4,000 tok/s and decode around 60 tok/s through the end of the 256K context.
Detailed benchmarks:
- Prompt lengths 10k–250k: TTFT scales linearly from 2.1s to 63s; ITL stays at 16.7–17.7ms. With MTP, ITL often drops to 12–16ms and output rises to 60–82 tok/s at some lengths
- Concurrency 1–16: at concurrency 1, MTP gives +35.7% output (55→75 tok/s); +9–12% at 4/8; at 16 it flips to -23%. MTP accept length 3.0–3.4
Repo: github.com/haohervchb/sglang-V100
More from Infra
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01