Alibaba’s Qwen-Image 3.0 aims at 4,500-token prompts and readable 10px text
matchaman11 · x · 2026-07-21
Alibaba’s Qwen-Image 3.0 is being presented as a major step up in image generation.
The post claims it can:
- Handle prompts up to 4,500 tokens
- Render tiny 10px text that remains readable
- Generate natively in 12 languages
- Create newspapers, exam papers, and complex UI layouts
- Produce full 3×3 infographic grids in one image
- Support 100+ visual styles
The examples in the image are meant to show unusually strong text rendering and layout control, especially for multilingual and document-style generation.
More from Multimodal
- TimeLens2 claims SOTA on 7 video grounding benchmarks with 4B and 8B models — _akhaliq · 2026-07-21
- AI-made 4-minute horror short ‘THE NOT KNOW’ lands as a shareable demo — gen_ericai · 2026-07-21
- SVG Generation Comparison: Leading AI Models Draw a Red Ferrari — Able-Line2683 · 2026-07-21
- Adding order metadata makes VLM error detection collapse, new benchmark shows — m_wulfmeier · 2026-07-21
- Claude AGI Agent starts paging itself in Slack with a no-heartbeat alarm — Sauers_ · 2026-07-21
- Gemini Omni is being called a video-editing leap on par with Nano Banana — CodeByPoonam · 2026-07-21