Pushing Qwen3.8-27B to 381 tps on a single RTX 3090

iamMess · reddit · 2026-08-21

The author shares updates on their hyper-optimized Qwen3.8-27B inference engine, achieving 382 tps for document reproduction on an RTX 3090. Key optimizations include:

The author provides a honest correction, noting that quantized KV caches reduce MTP acceptance rates (dropping from 2.56 to 2.38), resulting in a 2.13x decode tax at 112k context. This mode is optimized for RAG and coding assistants that frequently quote context, rather than general chat.

Related event: Qwen3.8-27B Inference Hits 381 TPS on RTX 3090(2 posts)→

Original post →

More from Infra

Infra channel →