Uncensored Qwen3-VL 32B GGUF Runs on Minimum 6.7GB VRAM
Every-Walrus · reddit · 2026-08-05
A developer has released uncensored GGUF files for the Qwen3-VL-32B-Instruct model on Hugging Face.
- Extreme Compression: By removing unused parts of the text encoder, the model size is significantly reduced, requiring as little as 6.7GB VRAM to run.
- Toolchain: It requires the author's forked ComfyUI-GGUF loader to support GGUF quantization for native ComfyUI models.
Related event: Uncensored and Quantized Version of Qwen3-VL-32B Released(3 posts)→
More from Models
- Alibaba's Qwen 3.8-Max Activates Only 95B of 2.4T Parameters — Div_pradeep · 2026-08-05
- Qwen's FinIndices Benchmark Exposes Severe Bottlenecks in LLM Financial Reasoning — Qwen · 2026-08-05
- User Notes Claude Opus Drops Pleasantries for Blunt Direct Answers — CtrlAltDwayne · 2026-08-05
- Abliteration: Removing LLM Safety Guardrails Without Retraining is Now an Open-Source Standard — maximelabonne · 2026-08-05
- DeepSeek V4 Flash Local Benchmark: MXFP4 Quantization Balances Speed and Top Scores — WonderRico · 2026-08-05
- OpenAI Codex Infinite Loop Bug Suspected of Doubling Token Usage — nlight · 2026-08-05