Running Qwen3.8-Flash-Next on 96GB VRAM: llama.cpp settings hit 15 t/s at 130k ctx

HlddenDreck · reddit · 2026-09-09

A Reddit user shares their full llama.cpp config for running Qwen3.8-Flash-Next (unsloth UD-Q4KXL GGUF) on 96GB VRAM: layer split-mode, flash-attn on, fit-ctx 262144, cache-ram 94208, and Qwen-style sampling (top-p 0.95, top-k 20). They report 15 t/s generation and 100-200 t/s prefill at 130k context, and are asking whether embeddings get offloaded to RAM and how others tune similar setups.

Original post →

More from Infra

Infra channel →