Optimal llama.cpp settings for Qwen 3.8 27B on RTX 6000 Pro

vhthc · reddit · 2026-08-25

A user shared llama-server configurations for running Qwen 3.8 27B (BF16) on a single RTX 6000 Pro with 256kb context. Key settings include draft-mtp speculative decoding for 2x speedup and disabling mmap to avoid ZFS issues. Current performance is 50-60 t/s generation. The author seeks advice on optimizing vllm or sglang for model swapping.

Original post →

More from Infra

Infra channel →