Squeezing Qwen3.8-27B With 256K Context Into 24GB: Mixed NVFP4 Quant Hits 50 tok/s
iam31337 · reddit · 2026-08-18
The author fit Qwen3.8-27B (embedded MTP included) with its full 262,144-token context onto a 24GB RTX PRO 4000 Blackwell SFF while keeping quality and speed:
- Memory: 23,952/24,467 MiB used (97.9%), 515 MiB left after 261,500 tokens; the F16 vision projector sits on a second GPU (982 MiB)
- Calibration: real production history over generic corpora — 5,472 messages from 296 Hermes agent sessions (coding, tool calls, infra, Polish/English mix)
- Tensor recipe: llama-imatrix ranked 497 target weights as a tensor map; bulk matrices in native NVFP4, sensitive attention/DeltaNet/FFN in Q5K/Q6K, embeddings Q6K, output head Q80, MTP layer NVFP4
- Results: 16.3GB GGUF at 5.01 BPW; WikiText-2 PPL 6.1197 vs 6.1127 for Q41 (0.11% gap); a stock NVFP4 quant hit 6.4949 — full FP4 was too aggressive
- Performance: 50.44 tok/s production average; target-only decode 21.19, MTP boosts to 59.46 (2.81x); custom llama.cpp 55.40 vs 45.42 master (+21.97%); at full 261.5K context decode drops to 12.61 tok/s (bandwidth-bound KV reads)
- Counterintuitive: a higher-precision MTP drafter made it 26.6% slower — the more accurate drafter matched the quantized target less often
GGUF published on HuggingFace; full tensor recipe, llama.cpp patches, runtime args and failed experiments on the author's blog.
More from Infra
- Guide: squeeze ~18-20 tok/s from Qwen3.8-27B on 16GB VRAM with lossless KV cache — BassAzayda · 2026-08-18
- Brex benchmark: Over half of top-growing startups build AI infra — mattturck · 2026-08-18
- Falcata: CUDA-native GBDT rebuilds training loop, 14x faster than LightGBM — srchvrs · 2026-08-18
- Brex Data: 14 of Top 25 Fastest-Growing Vendors Are AI Infrastructure — AccBalanced · 2026-08-18
- NVIDIA Boosts Local AI with Unsloth Integration and llama.cpp Optimizations — danielhanchen · 2026-08-18
- Text Watermark Detection Does Not Require Rerunning the LLM — rasbt · 2026-08-18