Running Qwen3.8-Flash-Next on 96GB VRAM: llama.cpp settings hit 15 t/s at 130k ctx
HlddenDreck · reddit · 2026-09-09
A Reddit user shares their full llama.cpp config for running Qwen3.8-Flash-Next (unsloth UD-Q4KXL GGUF) on 96GB VRAM: layer split-mode, flash-attn on, fit-ctx 262144, cache-ram 94208, and Qwen-style sampling (top-p 0.95, top-k 20). They report 15 t/s generation and 100-200 t/s prefill at 130k context, and are asking whether embeddings get offloaded to RAM and how others tune similar setups.
More from Infra
- Broadcom ASICs podcast: GPU supply means little when substrates and MLCCs bottleneck output — BenBajarin · 2026-09-09
- 27B 1-bit model runs in the browser at 25-30 tok/s on a 6GB RTX 3060 laptop (WebGPU) — mentria-ai · 2026-09-09
- The expensive part of coding agents isn't the agents—it's the silent retries ($900 for one task) — mrtrly · 2026-09-09
- AI infra is unbundling: model, harness, inference and compute become four separate choices — _changxu · 2026-09-09
- Running a 27B Q4 Model at 131K Context on a Single RTX 3090: ~25.6 tok/s Measured — bjivanovich · 2026-09-09
- Polymarket puts 73% odds on a US state enacting a data center moratorium by 2026 — Polymarket · 2026-09-09