Qwen3.8-27B uncensored quants released, FastMTP boosts inference up to 3.02x
hauhau901 · reddit · 2026-08-18
Community quant author HauhauCS released Qwen3.8-27B Uncensored Aggressive with the full KP quant range, vision support, and native NextN. The Aggressive uncensoring profile scored 0 refusals across 465 test prompts while keeping the base model's reasoning, agentic, and multimodal capabilities. Over 400 people requested access before launch, and the author's models are approaching 30M downloads on Hugging Face.
The headline feature is HauhauCS FastMTP: in Q8KP service tests it reached up to 3.02x document token generation and 1.93x reasoning TG versus MTP disabled, and 35.2%/21.1% more than the standard embedded MTP profile, with every drafted token still verified by the target model. A single 903 MB sidecar works across all quants, paired with a llama.cpp patch; exact build and serving commands are in the README.
Other notes:
- KP quants are model-specific profiles gaining 1–2 quant levels of quality for 5–15% more size
- 27B dense, 64 layers (48 Gated DeltaNet + 16 gated attention), 262K native context
- Full sampling params for thinking/non-thinking modes; author warns of malicious GGUF payloads distributed under similar names and provides checksums
More from Infra
- Grid Bottlenecks Stall AI: Interconnection Queues Surge to 45 Months — PeterDiamandis · 2026-08-18
- Cooling solutions for multi-3090 setup for local inference in 2026 — Sevealin_ · 2026-08-18
- Nvidia secures 35%-40% of global HBM supply for next year — JOBhakdi · 2026-08-18
- pagedMark: Invisible SynthID watermark removal optimized for Apple Silicon — d0ofz · 2026-08-18
- Ex-SpaceX engineers build AI robotic factory for steel parts — Ars Technica AI · 2026-08-18
- Benchmarking Qwen3.8-27B on 4x RTX 3090: Topology Matters — Mr_Moonsilver · 2026-08-18