Qwen 27B on RTX 5070 Ti laptop: 4.5 tok/s with 80% MTP acceptance

CoffeeToCode99 · reddit · 2026-08-15

Running Qwen3.8-27B (UD-Q4KXL) on a laptop with a 12GB RTX 5070 Ti GPU, utilizing CPU offloading and MTP speculative decoding, achieved approximately 4.5 tok/s with an MTP acceptance rate around 80%. Tests show stable performance for factual queries and Python coding, making it a viable setup prioritizing quality over speed.

Original post →

More from Infra

Infra channel →