How to run Qwen3.8-Flash-Next with N-gram SSD streaming in llama.cpp?

Ambitious_Fold_2874 · reddit · 2026-09-17

The author asks which llama-server flags to use for running Qwen3.8-Flash-Next with N-gram streaming off SSD in llama.cpp, and whether it's supported on the main branch or requires another one. Hardware: 4x 5060 Ti 16GB plus 256GB quad-channel DDR4.

Original post →

More from Infra

Infra channel →