Running Qwen3.8-Flash-Next on 2x5080: should I move from llama.cpp to vLLM?
whatyathinkk · reddit · 2026-09-17
A user runs Qwen3.8-Flash-Next Q5KXL on 2x RTX 5080 + 128GB DDR5 via llama.cpp with MTP speculative decoding and CPU MoE offload (ncmoe42), hitting 100 tps prefill and 18-20 tps decode. With few NVFP4/FP8 .gguf files available, they ask whether switching to vLLM/SGLang would yield better performance for this model on their setup. Full config included.
More from Infra
- OpenJev open-sources Jev-style semantic decisions on a single RTX 3090 with a frozen 4B model — alexcovo_eth · 2026-09-17
- MiniMax H3 video gen on dual 7900XTX takes 5.5min per 5s clip; user crowdsources Nvidia GPU numbers — HopefulConfidence0 · 2026-09-17
- Agent reliability: provider failover, circuit breakers, and idempotent retries for tool calls — Future_AGI · 2026-09-17
- Seagate IronWolf Pro HDDs doubled in a year, from $550 to $1,100 — HankYeomans · 2026-09-17
- Polymarket puts 31% odds on a US state enacting a data center moratorium by end of 2026 — Polymarket · 2026-09-17
- Data centers now supply roughly 45% of local tax revenue in Loudoun County, Virginia — Polymarket · 2026-09-17