Optimizing Qwen 3.8 27B with MTP and KV Cache Quantization
gabrielesilinic · reddit · 2026-08-31
Detailed guide on running Qwen 3.8 27B-UD-IQ4XS on a 7900XTX using llama-server.
Key Optimizations:
- MTP Tuning: Used --spec-draft-p-min 0.70 to reject low-confidence drafts, fixing slowdowns in long chats.
- KV Cache Quantization: Enabled q80 for both K and V caches without stability issues.
- Performance: Achieves 22-30 t/s at 140k context, utilizing 22.1GB VRAM.
Stack: llama-swap for management, systemd for service orchestration.
More from Infra
- Local Llama 3.1 install caused slowdown, fixed by uninstall — AiJohnAllen · 2026-08-31
- SimSlim tool fixes AI iOS coding crashes by running more simulators per Mac — bigblueboo · 2026-08-31
- Local 8B Model Document Extraction Demo on iPhone 16 — Better_Comment_7749 · 2026-08-31
- Apple May Scrap 2027 Mobile HBM Plans Due to High Costs — power97992 · 2026-08-31
- Matmul Energy Efficiency Competition Challenges AlphaTensor Algorithm — yaroslavvb · 2026-08-31
- NVIDIA DGX Station offers data-center-class performance — SpendLucky1273 · 2026-08-31