Config: Running Qwen3.8 27B Q8 on 3x RTX A4000

LevelSoft1165 · reddit · 2026-08-21

A user shared a config for running the Qwen3.8 27B Q8 quantized model on 3x RTX A4000 (16GB) GPUs using llama.cpp. With MTP optimization enabled, the setup achieves an average speed of about 23 tokens/s. The user is seeking advice for further performance improvements.

Original post →

More from Infra

Infra channel →