Optimal llama.cpp settings for Qwen 3.8 27B on RTX 6000 Pro
vhthc · reddit · 2026-08-25
A user shared llama-server configurations for running Qwen 3.8 27B (BF16) on a single RTX 6000 Pro with 256kb context. Key settings include draft-mtp speculative decoding for 2x speedup and disabling mmap to avoid ZFS issues. Current performance is 50-60 t/s generation. The author seeks advice on optimizing vllm or sglang for model swapping.
More from Infra
- UC Berkeley open-sources FreeToken: MoE inference 2-4x faster, runs 35B model on 8GB GPU — Roger_M_Taylor · 2026-08-25
- AI Chipmaker Enflame Sets Subscription Date for Near $900M Shanghai IPO — pstAsiatech · 2026-08-25
- Taiwan indicts 9 as 74 Nvidia B300 AI servers allegedly smuggled into China — kimmonismus · 2026-08-25
- Qwench: Fully automated local fine-tuning pipeline for Qwen3.8-27B — raiyanyahya · 2026-08-25
- Keploy MCP connector: Generate API tests from traffic or specs — modelcontextprotocol · 2026-08-25
- Nvidia Pays $6B for Poolside Tech and Staff, Not the Company — PrajwalTomar_ · 2026-08-25