Running Qwen 3.8 27B at 75t/s on 16GB VRAM: A Practical Guide

Kernoriordan · reddit · 2026-08-31

The author shares their experience tuning Qwen 3.8 27B on an RTX 4080/5080 (16GB VRAM), achieving an average decode speed of 75t/s (peaks at 100t/s) using a specific quantization and llama.cpp configuration.

Key Setup:

Logs demonstrate consistent performance during long text generation.

Original post →

More from Infra

Infra channel →