Unsloth GGUFs run Qwen-Image-2.1 locally on as little as 6GB VRAM
danielhanchen · x · 2026-09-23
Unsloth released Dynamic GGUF quantizations of Qwen-Image-2.1, claiming the 7B text-to-image model matches Nano Banana 2.0 and runs locally on 12GB VRAM — or 6GB with FP8 via RAM offloading.
- INT8/FP8 with pinned offloading fits in 6-8GB VRAM while staying reasonably fast
- The GGUF contains only the denoiser; pair with the bf16 VAE and Qwen3-VL-8B text encoder. Measured: UD-Q4KXL encoder vs Q4KM gives LPIPS 0.029, SSIM 0.959, 5.15GB vs 5.03GB, 36.5s vs 39.0s per image
- Runs via Unsloth Desktop or stable-diffusion.cpp; model card includes a full sd-cli command (1024x1024, 20 steps, cfg 6.0, euler)
- Qwen-Image-2.1 itself is a unified text-to-image and editing model with 32 Single-Stream DiT layers
Related event: Unsloth's quantized Qwen-Image-2.1 runs locally on just 6GB VRAM(2 posts)→
More from Infra
- 421M-Parameter Laya Model Plays Flappy Bird on CPU via OpenVINO INT8 — simpleuserhere · 2026-09-23
- How to run Qwen3.8-27B with 160k context on a 16GB AMD card: full config — According_Study_162 · 2026-09-23
- MLX-Serve v26.9.5 lands with Qwen-Image 2.1 and 4-way MTP streams at up to 122 tok/s on M4 Max — TheMoonMidas · 2026-09-23
- DigitalOcean Managed Agents enters public preview: idle pausing, 75+ models, one bill — HeyAmit_ · 2026-09-23
- Swapping just the decision layer: Qwen + SGLang beats Jev by ~37% at same accuracy — VeryWellVersed · 2026-09-23
- Reka EdgeQ VLM Runs Natively on Snapdragon 8 Elite NPU With 0.73s First Token — RekaAILabs · 2026-09-23