RTX 4090 hits 90 t/s on Qwen3 27B with llama.cpp — full config shared

SirApprehensive7573 · reddit · 2026-08-21

A user shares an exact llama.cpp setup running Qwen3-27B (UD-IQ4XS) on a fully-free RTX 4090: all GPU layers, FlashAttention, q80 KV cache, batch 256/ubatch 64, plus MTP-2 speculative decoding, reaching 90 t/s at 242k context. Without MTP the full 262k context fits but speed drops to 45-50 t/s; he pairs it with opencode on large codebases and prefers speed. Other Q4 quants and FP16 KV cache showed no noticeable difference for this model.

Original post →

More from Infra

Infra channel →