SenseTime Launches Multimodal Agent Base U1 Pro and Open-Source Vision Dataset

During WAIC 2026, SenseTime released SenseNova U1 Pro, a flagship native multimodal foundation model, and open-sourced the SenseNova-Vision dataset containing 50 million samples. The model aims to upgrade multimodal system capabilities from simply "providing answers" to directly "delivering finished products," attracting significant industry attention.

Core Technology and Product Positioning

According to reports from Synced and APPSO, U1 Pro is built on the NEO-Unity architecture, unifying understanding, generation, and action into a native interleaved vision-language reasoning capability. Unlike conventional entertainment-level image generation, the model focuses on native 8K output and "delivery-level" high-precision visual content generation. It is primarily designed for long-horizon delivery tasks, capable of handling professional design scenarios such as complex Chinese text, infographics, commercial posters, academic design, and industrial illustrations. Compared to the previous U1 Lite version, the official release specifically emphasized its enhanced ability to control image-text details.

Practical Applications and Review Feedback

In hands-on tests and reports by media outlets like APPSO, U1 Pro demonstrated the ability to generate complex visual content, including World Cup match reports and original movie posters. A article from Xin Zhi Yuan noted that this marks SenseTime's multimodal technology's practical potential to transition from conceptual understanding to the direct delivery of final visual products when tackling complex tasks.

2026-07-20 ~ 2026-07-21 · 6 related posts

Primary sources