Qwen3.8-Flash-Next at full context: 35s to first token on SGLang vs 258s on llama.cpp

FantasticNature7590 · reddit · 2026-09-09

A developer benchmarked Qwen3.8-Flash-Next across llama.cpp, SGLang and FreeToken on the same workstation (RTX PRO 6000 Blackwell 96GB, Ryzen 9 9950X, 96GB DDR5), including new PRs and speculative decoding.

Key numbers

Caveat: quantization (UD-IQ4XS GGUF vs NVFP4), KV-cache format, memory placement and speculative decoding differ across stacks (SGLang uses a NEXTN draft head; FreeToken ran without speculation), so results don't isolate engine software alone — full configs are in the repo. The author also honestly withdrew the PLE direct-read comparison after discovering a renamed flag was silently ignored.

Original post →

More from Infra

Infra channel →