RTX 4090 hits 90 t/s on Qwen3 27B with llama.cpp — full config shared
SirApprehensive7573 · reddit · 2026-08-21
A user shares an exact llama.cpp setup running Qwen3-27B (UD-IQ4XS) on a fully-free RTX 4090: all GPU layers, FlashAttention, q80 KV cache, batch 256/ubatch 64, plus MTP-2 speculative decoding, reaching 90 t/s at 242k context. Without MTP the full 262k context fits but speed drops to 45-50 t/s; he pairs it with opencode on large codebases and prefers speed. Other Q4 quants and FP16 KV cache showed no noticeable difference for this model.
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24