Qwen3.8-27B on 24GB VRAM: 131k Context with MTP Enabled
sisyphus-cycle · reddit · 2026-08-17
A user shared their llama.cpp configuration for running Qwen3.8-27B (Q4KM) on a 24GB VRAM GPU (RTX 4090). The testing reveals a trade-off between context length and generation speed controlled by MTP (Multi-Token Prediction). Without MTP, the model achieves 194k context tokens at 40 tps. With MTP enabled, context drops to 131k tokens, but generation speed increases to 65 tps (peaking at 80 tps for Python code).
More from Infra
- Enterprise AI privacy risks highlight the need for sovereign AI architectures — Familiar-Display2989 · 2026-08-17
- TensorCast Unifies Tensor Management, Boosting LLM Startup Speed by 228x — 机器之心 · 2026-08-17
- Running dstack Confidential VMs for private code execution on cloud — bgmshana · 2026-08-17
- AI for hardware engineering: Can models understand and improve complex structures? — rms80 · 2026-08-17
- US per capita power consumption peaked at dot-com,暗示 scaling limits — jwt0625 · 2026-08-17
- User runs Krea2 and MiniMax Music locally — -becausereasons- · 2026-08-17