Pushing Qwen3.8-27B to 124 tps on a single RTX 3090

iamMess · reddit · 2026-08-19

A developer demonstrated pushing Qwen3.8-27B inference to 124 tps (greedy) on a single RTX 3090 through extreme optimization. Key techniques include:

All optimizations maintain exact output distribution. The code is open-sourced.

Related event: Qwen3.8-27B Inference Optimized for RTX 3090(2 posts)→

Original post →

More from Infra

Infra channel →