How to run Qwen3.8-Flash-Next with N-gram SSD streaming in llama.cpp?
Ambitious_Fold_2874 · reddit · 2026-09-17
The author asks which llama-server flags to use for running Qwen3.8-Flash-Next with N-gram streaming off SSD in llama.cpp, and whether it's supported on the main branch or requires another one. Hardware: 4x 5060 Ti 16GB plus 256GB quad-channel DDR4.
More from Infra
- ROCm Beats Vulkan on Strix Halo: Up to 47% Faster Decoding at 41.9 tok/s — QuixiAI · 2026-09-17
- Qwen3.8-Flash on 4x AMD V620 hits 1,300 PP and 70+ tok/s on coding — Thin_Pollution8843 · 2026-09-17
- Retraining fp16 scales only: patching a 3-bit Qwen3.8-27B GGUF closer to its BF16 parent — ZenZombie117 · 2026-09-17
- New dashboard sets first public baseline for advanced RL training costs — teortaxesTex · 2026-09-17
- PyTorch Conference NA lineup spotlights torch.compile and custom kernel breakthroughs — PyTorch · 2026-09-17
- Structured-decision trick speeds up DiffusionGemma inference 3-10x with one forward per request — bodonoghue85 · 2026-09-17