Running Qwen 3.6 27B on RTX 5090: 40 t/s at 262k Context in llama.cpp

Gargle-Loaf-Spunk · reddit · 2026-08-08

A developer shared their optimized llama.cpp launch arguments for running the Qwen 3.6 27B model (Q6K quantization) on an RTX 5090, specifically tailored for app development tasks.

Performance

Key Parameter Tuning

The setup also enables advanced features like MTP speculative sampling (--spec-type draft-mtp) to maximize performance.

Original post →

More from coding & agent

coding & agent channel →