Qwen3.8-27B uncensored quants released, FastMTP boosts inference up to 3.02x

hauhau901 · reddit · 2026-08-18

Community quant author HauhauCS released Qwen3.8-27B Uncensored Aggressive with the full KP quant range, vision support, and native NextN. The Aggressive uncensoring profile scored 0 refusals across 465 test prompts while keeping the base model's reasoning, agentic, and multimodal capabilities. Over 400 people requested access before launch, and the author's models are approaching 30M downloads on Hugging Face.

The headline feature is HauhauCS FastMTP: in Q8KP service tests it reached up to 3.02x document token generation and 1.93x reasoning TG versus MTP disabled, and 35.2%/21.1% more than the standard embedded MTP profile, with every drafted token still verified by the target model. A single 903 MB sidecar works across all quants, paired with a llama.cpp patch; exact build and serving commands are in the README.

Other notes:

Original post →

More from Infra

Infra channel →