Qwen 27B on RTX 5070 Ti laptop: 4.5 tok/s with 80% MTP acceptance
CoffeeToCode99 · reddit · 2026-08-15
Running Qwen3.8-27B (UD-Q4KXL) on a laptop with a 12GB RTX 5070 Ti GPU, utilizing CPU offloading and MTP speculative decoding, achieved approximately 4.5 tok/s with an MTP acceptance rate around 80%. Tests show stable performance for factual queries and Python coding, making it a viable setup prioritizing quality over speed.
More from Infra
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24
- Hyperscalers: Choosing Between HDD and SSD Based on Space and Cost — generativist · 2026-08-24
- Samsung shows new HBM cooling solution, hints at die performance variance — BenBajarin · 2026-08-24
- Tobi open-sources walgit: A single-binary Git server backed by object stores — jevon · 2026-08-24
- s3collections: Durable Go data structures backed directly by S3-compatible storage — andersonbcdefg · 2026-08-24
- Prediction market gives 68% chance of a state data center moratorium by year-end — Polymarket · 2026-08-24