Pushing Qwen3.8-27B to 381 tps on a single RTX 3090
iamMess · reddit · 2026-08-21
The author shares updates on their hyper-optimized Qwen3.8-27B inference engine, achieving 382 tps for document reproduction on an RTX 3090. Key optimizations include:
- Longer verify blocks: Increasing verification to 16 tokens per step using a lookup drafter, achieving a 15/16 acceptance rate in document-quoting workloads.
- Breaking the 64K context limit: Using int8 KV cache and adjusting sliding window block sizes to expand context support from 69k to 138k tokens.
The author provides a honest correction, noting that quantized KV caches reduce MTP acceptance rates (dropping from 2.56 to 2.38), resulting in a 2.13x decode tax at 112k context. This mode is optimized for RAG and coding assistants that frequently quote context, rather than general chat.
Related event: Qwen3.8-27B Inference Hits 381 TPS on RTX 3090(2 posts)→
More from Infra
- AWS built an MCP server for 16,000 APIs, discussing agent sprawl and minimalist architecture — dsp_ · 2026-08-21
- Why Bittensor ($TAO) Could Be the Next Bitcoin or Ethereum: A Deep Dive into Tokenomics — bittingthembits · 2026-08-21
- Employees Connecting AI Tools Internally Risks Leaks; Merge API Adds DLP — shensi · 2026-08-21
- Private Clouds Offer Control Over Hyperscale 'Noisy Neighbor' Issues for AI — DavidLinthicum · 2026-08-21
- CoreWeave Wins Multi-Billion Dollar AI Cloud Deal with Hudson River Trading — _ScottCondron · 2026-08-21
- Polymarket: 14% chance AI bubble bursts by year-end — Polymarket · 2026-08-21