Squeezing Qwen3.8-27B With 256K Context Into 24GB: Mixed NVFP4 Quant Hits 50 tok/s

iam31337 · reddit · 2026-08-18

The author fit Qwen3.8-27B (embedded MTP included) with its full 262,144-token context onto a 24GB RTX PRO 4000 Blackwell SFF while keeping quality and speed:

GGUF published on HuggingFace; full tensor recipe, llama.cpp patches, runtime args and failed experiments on the author's blog.

Original post →

More from Infra

Infra channel →