Testing Qwen3.8-27B with 131k Context on 16GB VRAM using Speculative Decoding
BuffMcBigHuge · reddit · 2026-08-20
The author tested Qwen3.8-27B on an RTX 4080 16GB to maximize context and performance. By using a custom llama.cpp build with the DFlash2 branch, ngram-mod sampling, and MTP (Medusa Thought Proposal) speculative decoding, they achieved stable context lengths up to 131k. The post compares results across different configurations like IQ3XXS quantization and provides full startup commands.
More from Infra
- New Disaggregation for Hybrid Linear Models on Cerebras CS-4 — AccBalanced · 2026-08-20
- Cursor Deep Dive: 20 Years of Git Infrastructure Evolution Lead to Database-Like Storage Design — omojumiller · 2026-08-20
- NVIDIA DGX Station serves Qwen3.8-27B at 2,713+ tok/s peak — NVIDIAAI · 2026-08-20
- ECB blog predicts AI market correction is coming — marigo · 2026-08-20
- S&P data predicts top 5 tech firms could burn $125B cash by 2027 — marigo · 2026-08-20
- Is 5t/s KIMI K3 on a Single 5090 Feasible? — MLDataScientist · 2026-08-20