Running Qwen3.8-Flash-Next on 2x5080: should I move from llama.cpp to vLLM?

whatyathinkk · reddit · 2026-09-17

A user runs Qwen3.8-Flash-Next Q5KXL on 2x RTX 5080 + 128GB DDR5 via llama.cpp with MTP speculative decoding and CPU MoE offload (ncmoe42), hitting 100 tps prefill and 18-20 tps decode. With few NVFP4/FP8 .gguf files available, they ask whether switching to vLLM/SGLang would yield better performance for this model on their setup. Full config included.

Original post →

More from Infra

Infra channel →