Qwen3.8-Flash-Next NVFP4 runs full 256K context on 4x V100 in SGLang-V100

Primary_Exchange21 · reddit · 2026-08-30

The community project SGLang-V100 now supports RadixArk/Qwen3.8-Flash-Next-NVFP4, running full context on 4x V100 32GB with >50GB of ngram offloaded to system RAM. Prefill holds around 4,000 tok/s and decode around 60 tok/s through the end of the 256K context.

Detailed benchmarks:

Repo: github.com/haohervchb/sglang-V100

Original post →

More from Infra

Infra channel →