Qwen3.8-27B hits 134 TPS on RTX 3090 with DFlash2 and custom optimizations

iamMess · reddit · 2026-08-20

The author pushed Qwen3.8-27B inference on an RTX 3090 to 134 TPS (default sampling) and drastically reduced long-context turn latency from 23s to 1s.

Key Optimizations:

Performance:

Related event: Qwen3.8-27B Inference Hits 381 TPS on RTX 3090(2 posts)→

Original post →

More from Infra

Infra channel →