Qwen 3.8 27B hits 96 t/s decode with 110k context on a single 16GB RTX 5080 via NInfer
Kernoriordan · reddit · 2026-10-07
A developer improved on their earlier 75 t/s llama.cpp setup by running Qwen 3.8 27B with NInfer v1.5 on a 16GB RTX 5080 (Ubuntu 24.04 via WSL2), achieving 90–110 t/s decode with 110,592 tokens of context.
Real-world numbers (32 Zoo Code coding requests)
- Median decode 96.45 t/s (range 84.6–131.7)
- Median time to first token 1.4s, with prompt caching hitting on most turns
- Prompts grew to 77–79k tokens in the session
Key config
- Weights take 11.86 GiB, leaving 498 MiB slack; prefill chunk reduced to 896 to fit
- q4 KV cache, MTP speculative decoding (draft 3, 58–67% acceptance), 2048 thinking budget
- 68% of generated tokens were reasoning; a cache miss on a 79k prompt cost 50s
Gotchas
- Zoo Code's base URL must include /v1 or you get a 404
- Zoo Code sends high reasoning effort despite showing medium; disable it client-side
No controlled quality comparison vs GGUF yet.
More from Infra
- One RTX Pro 6K Sidecar Lifts DGX Station GB300 to 60K tok/s Prefill with DeepSeek V4.1 Flash — Sentdex · 2026-10-07
- Hugging Face's Talk: A Full Tour of llama.cpp and the Local AI Inference Ecosystem — unofficialmerve · 2026-10-07
- Constellation Energy stock surges 14% on 20-year nuclear power deal with Google — Polymarket · 2026-10-07
- DNS root KSK rollover hits October 11: Cloudflare explains KSK-2024 switch and readiness test — Cloudflare Blog · 2026-10-07
- Qwen3.8 Flash Next IQ1_M hits 55 tok/s on a 5060 Ti 16GB and still codes well — bobaburger · 2026-10-07
- OpenAI to fund Nicholas Nethercote's work speeding up the Rust compiler — charliermarsh · 2026-10-07