SenseTime Releases U1Pro Multimodal Creation Model

新智元 · wechat · 2026-07-18

At WAIC, SenseTime released its next-generation multimodal model, **SenseNova U1Pro**, featuring a unified base for "understanding, generation, and action." The article primarily showcases its image generation capabilities in tasks like **8K scrolls, movie posters, infographics, Chinese long-text typography, and character design sheets**, emphasizing that it acts more like a "thinking designer" rather than a slot-machine-style image generator. Technically, U1Pro employs a proprietary unified architecture called **NEO-unify**, integrating understanding, generation, and action into the same representation and Transformer sequence. The creation process is completed through long-range iterations similar to an **Agentic Generation Loop**. The article also notes: - Training incorporates a "Chain-of-Thought instruction" and multi-dimensional rewards to enhance composition, text accuracy, design aesthetics, and consistency - By utilizing larger patches and hierarchical noise control, it reduces the context and compute pressure of direct **8K output** - The company invited 200 art academy students and designers for review, claiming that currently, only U1Pro and GPT-Image-2 meet the threshold for being "directly deliverable to clients" - Future goals extend beyond visual creation to spatial intelligence, embodied brains, urban planning, and architectural design

Related event: SenseTime Launches SenseNova U1 Pro Native 8K Multimodal Model(6 posts)→

Original post →

More from Models

Models channel →