Qwen 3.8 27B quantization achieves 200k context on 16GB VRAM
abskvrm · reddit · 2026-08-28
A user achieved over 200k context length with the Qwen 3.8 27B UD-IQ3XXS model on 16GB VRAM. Compared to the previous UD-Q3KXL setup, prompt processing speed dropped from 700-800 tk/s to 400 tk/s, while quality differences are yet to be fully tested. KV cache was quantized to q51.
Related event: Quantized Qwen 3.8 27B Runs 200K Context on Low VRAM(2 posts)→
More from Infra
- Puro-2B Matches Qwen2.5 Performance with $6.9K Pretraining on RTX 5090 — _reachsumit · 2026-08-28
- Intel XE3P projected specs: 1.3 PFLOPS FP8, 1.5TB/s bandwidth, 2027 launch — QuixiAI · 2026-08-28
- Rumor: Anthropic interested in developing its own training chip — zephyr_z9 · 2026-08-28
- India Commits $13.4B for 'Semicon 2.0' Chip Design and Manufacturing — SumitGup · 2026-08-28
- Meta, Google, NTT to discuss AI data center optical architectures — jwt0625 · 2026-08-28
- Coherent-lite Catches Up to IMDD in Energy Efficiency for Pluggables — jwt0625 · 2026-08-28