Community finetunes Qwen-Image-2.1 VAE to kill checkerboard artifacts, plus realtime AE
bdsqlsz · x · 2026-09-30
Developer madebyollin released two unofficial VAE variants for Qwen-Image-2.1:
- Texture-Fix-VAE: the original VAE decoder finetuned for 5000 steps at lr 3e-5, with only the two highest-resolution decoder stages and output head unfrozen (7.5M trainable params). It removes checkerboard artifacts, with the biggest gains on detailed photo-style images.
- A distilled tiny AE for realtime encode/decode, e.g. live previews.
Both are on Hugging Face with ComfyUI (drop into models/vae and swap in the Load VAE node) and Diffusers integration snippets included.
More from Multimodal
- Argil soft-launches AI storytelling studio, cutting 30s video production from 3 hours to 10 minutes — kevinnbass · 2026-09-30
- Plenio 0.3 ships: local AI music production in ComfyUI with score editor, stems and mastering — Vivid_Promise1700 · 2026-09-30
- NVIDIA's Physis-Lang puts physics reasoning in captions, tops Physics-IQ with Cosmos 3 — NVIDIAAI · 2026-09-30
- One image + one mp3: Opus 5.5 renders a 20-second video scene in 30 minutes — justin_hart · 2026-09-30
- This AI video is so realistic the author's grandma asked if it was AI — techhalla · 2026-09-30
- ComfyUI seeks testers for new Asset System: indexed files, persistent output history — Lexius2129 · 2026-09-30