Qwen3.8-Flash-Next on 6x3090 without NVLink: prefill 8-10x faster, long-context decode 2-3x
flynth92 · reddit · 2026-10-08
A dev patched llama.cpp for multi-GPU long-context serving of Qwen3.8-Flash-Next over plain PCIe (6x3090/6x4090): prefill at 250k goes 8-10x faster, concurrent long-context decode up to 10.8x, with flat performance from 5k to 250k vs upstream halving every 50k. Binaries, docker images and per-patch commits are public; n-gram speculation adds 2-2.5x on code tasks. Caveats: one model tested, CUDA only, LLM-written code validated by measurement.
More from Infra
- Surface Laptop Ultra launches Oct 16 starting at $2,599 with Nvidia RTX Spark — tomwarren · 2026-10-08
- Microsoft Surface AI devices ship Friday: RTX Spark Surface Ultra from $2,599, Dev Box $5,999 — ryanshrout · 2026-10-08
- Microsoft's Surface RTX Spark Dev Box opens at $5,999 with 128GB unified memory — The Verge AI · 2026-10-08
- XPU Grasshopper claims AI-co-designed chip, 816x faster in 13 weeks — ycombinator · 2026-10-08
- NSA spending billions of dollars a year testing frontier AI models, sources say — coherence · 2026-10-08
- Hyperscalers will do anything to shave a microcent off pluggable transceiver costs — jwt0625 · 2026-10-08