Qwen3.5 27B GGUF Quantization Launched for 12GB VRAM

soyaakinohara · hf · 2026-08-23

A new quantized model of the Qwen3.5 series, soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf, is trending on Hugging Face. Optimized for llama.cpp, this version uses the GGUF format. The quantization reduces VRAM requirements to around 12GB, making it suitable for local deployment on consumer hardware. The model is tagged as "abliterated" and "uncensored", utilizing hybrid-attention and MTP (Multi-Token Prediction) techniques.

Original post →

More from Infra

Infra channel →