Qwen3.8-27B optimization hits 1150 tps on RTX 3090

iamMess · reddit · 2026-08-18

The author released an update to their hyper-optimized Qwen3.8-27B inference engine for RTX 3090. By implementing fp16 recurrent state, int8 activations across all layers, and probabilistic sampling, the engine achieves 99 tps for single requests and a peak of 1150 tps for batched requests. Prefill speeds increased by 25-50%. The project is open-sourced on GitHub.

Original post →

More from Infra

Infra channel →