Qwen-Image-2.1 specs leak: 7B DiT, Qwen3-8B encoder, 56GB VRAM for 4K
bdsqlsz · x · 2026-09-18
ModelScope opened 50 early-access spots for Qwen-Image-2.1, requiring participants to publish an original hands-on review by Sep 28. Developer bdsqlsz shared leaked specs: a 7B DiT (14GB), a Qwen3-8B text encoder, 1GB VAE, and 56GB VRAM needed for 4K (2048×2048) resolution.
More from Multimodal
- Prompt share: 'Digital Fracture' glitch-art template for image generation — azed_ai · 2026-09-18
- LightOnOCR-2-1B: 1B OCR model beats rivals 9x its size, 493K pages/day on one GPU — thisguyknowsai · 2026-09-18
- 2D animations in 15s: full GPT spritesheet + H3 Max Magnific workflow shared — techhalla · 2026-09-18
- KITScenes Multimodal Dataset Launches to Fill the Data Gap for 3D Foundation Models — abursuc · 2026-09-18
- Recreating videos in MiniMax H3: a 4-step frame-extraction workflow — Negative-Whereas3307 · 2026-09-18
- AI artist 'Will' drops a new song daily, with persistent memory and personality — kun101 · 2026-09-18