Qwen3.8-27B hits 268 tok/s on single Blackwell GPU with full context

EAccelerate_42 · x · 2026-08-24

A developer released an optimization recipe for Qwen3.8-27B on a single RTX PRO 5000 Blackwell (48GB). Using SGLang, DFlash2 block-16, and Triton attention, it achieves 267.8 tok/s (310.7 peak) with the full 262K context, doubling the EAGLE baseline. Reproducible scripts are provided.

Original post →

More from coding & agent

coding & agent channel →