Testing Qwen3.8-27B with 131k Context on 16GB VRAM using Speculative Decoding

BuffMcBigHuge · reddit · 2026-08-20

The author tested Qwen3.8-27B on an RTX 4080 16GB to maximize context and performance. By using a custom llama.cpp build with the DFlash2 branch, ngram-mod sampling, and MTP (Medusa Thought Proposal) speculative decoding, they achieved stable context lengths up to 131k. The post compares results across different configurations like IQ3XXS quantization and provides full startup commands.

Original post →

More from Infra

Infra channel →