Qwen3.8-Flash-Next on 6x3090 without NVLink: prefill 8-10x faster, long-context decode 2-3x

flynth92 · reddit · 2026-10-08

A dev patched llama.cpp for multi-GPU long-context serving of Qwen3.8-Flash-Next over plain PCIe (6x3090/6x4090): prefill at 250k goes 8-10x faster, concurrent long-context decode up to 10.8x, with flat performance from 5k to 250k vs upstream halving every 50k. Binaries, docker images and per-patch commits are public; n-gram speculation adds 2-2.5x on code tasks. Caveats: one model tested, CUDA only, LLM-written code validated by measurement.

Original post →

More from Infra

Infra channel →