RTX 3090 gets 35 t/s on Qwen 3.8 27B

cviperr33 · reddit · 2026-08-15

A user reports running Qwen 3.8 27B (IQ4 NL quant) on an RTX 3090 via llama.cpp at 34-36 tokens/s. This is slower than the 50-60 t/s they recalled getting on v3.6, without speculative decoding enabled. Despite the speed dip, they feel the model quality is truly next-gen.

Original post →

More from Infra

Infra channel →