Running Qwen3.8-27B with 128K context on RTX 4000 Blackwell
mmkaywhatevers · reddit · 2026-08-29
A developer successfully ran Qwen3.8-27B on a 24GB RTX PRO 4000 Blackwell. By specifying the correct CUDA 13.3 path and enabling MTP3 speculative decoding with INT8 KV cache, they achieved 128K context support. Tests show MTP3 improved decode speed from 31 tok/s to 60 tok/s (1.9x gain), with manageable VRAM usage even at high context lengths.
More from Infra
- Stripe employee calls for agent-friendly internet infrastructure — jeff_weinstein · 2026-08-29
- Cognition hits $900M ARR but may burn $800M on Nvidia servers — thedealdirector · 2026-08-29
- Cisco Partners with Super Micro for Liquid-Cooled AI Racks — Beth_Kindig · 2026-08-29
- Neocloud Lambda Secures $1B Debt to Buy Nvidia Chips — TechCrunch AI · 2026-08-29
- Kimi K3 Full Fine-Tuning Live: GPU Requirements Drop by 40% — ypatil125 · 2026-08-29
- Audit reveals 64 GGUF quants mislabeled across 25 repos — Daxfortuna · 2026-08-29