True Q4 Qwen 27B at 13 tok/s and 61K Context on a 16GB RTX 5080
nofuture09 · reddit · 2026-09-02
A detailed first-hand writeup of running Unsloth's Qwen3.8-27B UD-Q4KM (16.46GB) on an RTX 5080 16GB with full 65,536 context. Key move: selective FFN offload — moving the 16 largest FFN tensor groups (2.764 GiB) to CPU while keeping attention/KV on GPU — yields stable 13.2 tok/s at 50-61K context, versus 6.6 tok/s with whole-layer offload in LM Studio. Surprise finding: MTP speculative decoding hurt throughput (8.6/7.8 tok/s), likely due to RAM bandwidth contention. All configs, benchmarks, and failed profiles are open-sourced on GitHub.
More from coding & agent
- Anthropic ships official Claude Fable 5.1 prompting guide with 16 fixes for agent pain points — xiaohu · 2026-09-02
- Claude Fable 5.1 prompting guide: 16 official tips and an agent migration checklist — xiaohu · 2026-09-02
- Polar Analytics launches Polar Operator, an AI operator for commerce teams in Slack — JaynitMakwana · 2026-09-02
- $40K of guitar gear won't make you Mike Rutherford — why coding skill still wins with AI agents — JFPuget · 2026-09-02
- Agents-run SEO pipeline: 8 draft PRs in two weeks, most runs under $1 — Pitiful-Surround-285 · 2026-09-02
- Vibe coding beats product hunting: build the exact tool you need — Philmod · 2026-09-02