True Q4 Qwen 27B at 13 tok/s and 61K Context on a 16GB RTX 5080

nofuture09 · reddit · 2026-09-02

A detailed first-hand writeup of running Unsloth's Qwen3.8-27B UD-Q4KM (16.46GB) on an RTX 5080 16GB with full 65,536 context. Key move: selective FFN offload — moving the 16 largest FFN tensor groups (2.764 GiB) to CPU while keeping attention/KV on GPU — yields stable 13.2 tok/s at 50-61K context, versus 6.6 tok/s with whole-layer offload in LM Studio. Surprise finding: MTP speculative decoding hurt throughput (8.6/7.8 tok/s), likely due to RAM bandwidth contention. All configs, benchmarks, and failed profiles are open-sourced on GitHub.

Original post →

More from coding & agent

coding & agent channel →