Qwen-Image-3.0 brings Alibaba’s image generation to dense text layouts and multilingual scenes
智东西 · wechat · 2026-07-21
- A Chinese media report says Qwen-Image-3.0 has arrived as Alibaba Qwen’s third-generation image model.
- The biggest upgrade is turning image generation into a practical production tool: it can handle up to 4.5k tokens of input and generate dense layouts such as newspapers, storyboards, and exams.
- The article says it can render 10 px small text, reproduce fine details like pores and hair, and natively support 12 languages.
- It also highlights stronger performance in text alignment, multi-panel narrative consistency, complex composition, and UI-like scene generation.
- The report says API access is opening on Alibaba Cloud Bailian and the Qwen AI platform, with QwenStudio and the Qwen app coming soon.
More from Multimodal
- Code-driven project syncs music, instruments and video into The Loom — Sauers_ · 2026-07-21
- Fable made the music, the instrument, and the video for The Loom — Sauers_ · 2026-07-21
- Open-source TTS list sorts models by license before quality for commercial shipping — mahimairaja · 2026-07-21
- AI pastiche turns Chris Cornell’s “American Nightmare” into a meme poster — 77sevens · 2026-07-21
- AI gaming demos now generate real-time sound to match world-model video — mark_k · 2026-07-21
- Krea2 users find a 4-step Raw plus 4-step Turbo workflow that preserves quality — PropagandaOfTheDude · 2026-07-21