Config: Running Qwen3.8 27B Q8 on 3x RTX A4000
LevelSoft1165 · reddit · 2026-08-21
A user shared a config for running the Qwen3.8 27B Q8 quantized model on 3x RTX A4000 (16GB) GPUs using llama.cpp. With MTP optimization enabled, the setup achieves an average speed of about 23 tokens/s. The user is seeking advice for further performance improvements.
More from Infra
- AI Data Centers Bypass Grid Delays with On-Site Gas and Solar — shensi · 2026-08-21
- Unsloth Dynamic V3 GGUFs: Q3 Outperforms Larger Q4 Models — danielhanchen · 2026-08-21
- VC Trend: Invest in Own GPU Clusters to Offer Compute at Cost to Portfolio Companies — beffjezos · 2026-08-21
- Heterogeneous Systems Outperform Frontier Models in New Benchmarks — ShahabBakht · 2026-08-21
- AMD MI300x vs NVIDIA H100: Real-world agent coding benchmark — locker73 · 2026-08-21
- Quit Claude Pro: running Qwen 3.8-27B locally on a 5090 to replace Claude Code — SOC_FreeDiver · 2026-08-21