Qwen3.5 27B GGUF Quantization Launched for 12GB VRAM
soyaakinohara · hf · 2026-08-23
A new quantized model of the Qwen3.5 series, soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf, is trending on Hugging Face. Optimized for llama.cpp, this version uses the GGUF format. The quantization reduces VRAM requirements to around 12GB, making it suitable for local deployment on consumer hardware. The model is tagged as "abliterated" and "uncensored", utilizing hybrid-attention and MTP (Multi-Token Prediction) techniques.
More from Infra
- Deep dive into leading AI chip architectures and software stacks — ycombinator · 2026-08-23
- Bloomberg Podcast: How the AI Industry Missed the Data Center Backlash Entirely — AccBalanced · 2026-08-23
- Hot Chips 2026 opens tomorrow at Stanford, focusing on AI infra and agents — ai · 2026-08-23
- Implementing crash-resilient long-running AI agent loops in Go — Soft_Flower5258 · 2026-08-23
- Ireland reportedly considers nuclear power to meet data center demand — Polymarket · 2026-08-23
- Dual 3090 Qwen 27B Full Context Setup: 3200 tok/min Achieved — CryptographerLow7817 · 2026-08-23