Alibaba launches Qwen-Image-3.0 with 4.5k-token prompts and 12-language rendering
Scobleizer · x · 2026-07-21
Alibaba has released Qwen-Image-3.0, the third-generation image generation model in the Qwen-Image line.
The company says the model is built around a single idea: “Real”.
- It supports up to 4.5k token inputs and can generate complex layouts such as newspapers, storyboards, and exam papers.
- It can render text as small as 10px and reproduce fine details such as pores and hair strands.
- It supports native rendering in 12 languages and can simulate common interfaces such as web pages, games, and livestreams.
- Alibaba positions it as a practical productivity tool, not just a “good-looking” image model.
More from Multimodal
- fal and Sequoia Host the World's First Generative Video Hackathon — Jamez_Chi · 2026-07-23
- A 3D blockout plus phone walkthrough becomes a video-generation workflow — Bart_Bialobrzeski · 2026-07-23
- Clean Plate LoRA examples show a practical video-cleanup workflow — nazihater3000 · 2026-07-23
- Fish Audio S2 fork jumps from 1.2 to 23 t/s on Apple Silicon — fogonthebarrow-downs · 2026-07-23
- Qwen-image 3.0 is now available on Runware — aziz4ai · 2026-07-23
- Reddit user shares a moody Wan 2.2 video made in ComfyUI — Independent-Ebb7658 · 2026-07-23