Qwen3.8-27B-int4-AutoRound (18GB) with working MTP spec decode
BusinessMud9586 · reddit · 2026-08-16
Release of the Qwen3.8-27B-int4-AutoRound quantized model, sized at 18GB, featuring support for working MTP spec decode.
More from Models
- llama.cpp integrates Dots3 Note model, scoring 78.4 on SWE-bench Verified — victormustar · 2026-08-16
- Post-training boosts GLM-5.3 to rival larger models, showing size isn't the bottleneck — haider1 · 2026-08-16
- Commits suggest Qwen 35B model removed, likely not releasing — Local-Cardiologist-5 · 2026-08-16
- Blind Test: Anime Girl 3D Scene Generation Across Models — Jeanodel · 2026-08-16
- 1-bit Quantized Qwen 27B Runs on 12GB VRAM at 92 tokens/s — zyxciss · 2026-08-16
- BlueBench Cyber Benchmark: GPT-5.6 Sol Leads, Kimi K3 Tops Open-Weights — vijaybolina · 2026-08-16