Qwen 3.8 27B quantization achieves 200k context on 16GB VRAM

abskvrm · reddit · 2026-08-28

A user achieved over 200k context length with the Qwen 3.8 27B UD-IQ3XXS model on 16GB VRAM. Compared to the previous UD-Q3KXL setup, prompt processing speed dropped from 700-800 tk/s to 400 tk/s, while quality differences are yet to be fully tested. KV cache was quantized to q51.

Related event: Quantized Qwen 3.8 27B Runs 200K Context on Low VRAM(2 posts)→

Original post →

More from Infra

Infra channel →