Running Qwen3.8-27B with 128K context on RTX 4000 Blackwell

mmkaywhatevers · reddit · 2026-08-29

A developer successfully ran Qwen3.8-27B on a 24GB RTX PRO 4000 Blackwell. By specifying the correct CUDA 13.3 path and enabling MTP3 speculative decoding with INT8 KV cache, they achieved 128K context support. Tests show MTP3 improved decode speed from 31 tok/s to 60 tok/s (1.9x gain), with manageable VRAM usage even at high context lengths.

Original post →

More from Infra

Infra channel →