Qwen3.8-27B optimization hits 1150 tps on RTX 3090
iamMess · reddit · 2026-08-18
The author released an update to their hyper-optimized Qwen3.8-27B inference engine for RTX 3090. By implementing fp16 recurrent state, int8 activations across all layers, and probabilistic sampling, the engine achieves 99 tps for single requests and a peak of 1150 tps for batched requests. Prefill speeds increased by 25-50%. The project is open-sourced on GitHub.
More from Infra
- Qwen3.8-27B uncensored quants released, FastMTP boosts inference up to 3.02x — hauhau901 · 2026-08-18
- AI Accelerator Shipments Forecast to Reach 16.3M in 2026, Up 62% — Beth_Kindig · 2026-08-18
- Ollama benchmarks: DeepSeek V3 Flash leads, Qwen wins quality but 30x slower — ollama · 2026-08-18
- Agent Boom Pushes Frontier Model Gross Margins to Over 85% — ben_j_todd · 2026-08-18
- Bittensor co-founder: building open, permissionless AI you can mine like Bitcoin — markjeffrey · 2026-08-18
- SGLang Reserves 18.5GB for GDN State, vLLM Doesn't: 5x KV Cache Gap on Qwen3.8-27B — SomeRandomGuuuuuuy · 2026-08-18