SenseTime Launches Multimodal Agent Base U1 Pro and Open-Source Vision Dataset

During WAIC 2026, SenseTime launched the multimodal agent base model "SenseNova U1 Pro" and open-sourced SenseNova-Vision, a visual dataset containing 50 million samples. The model aims to elevate multimodal systems from simply "providing answers" to directly "delivering finished products," drawing significant industry attention.

Core Technology and Product Positioning

Built on the NEO-Unity architecture, U1 Pro unifies understanding, generation, and action into a native interleaved image-text reasoning capability. Unlike conventional entertainment-grade image generation, this model focuses on native 8K output and "delivery-grade" high-precision visual content generation. It is primarily designed for long-horizon delivery tasks, capable of handling complex Chinese text, infographics, commercial posters, academic layouts, and industrial diagrams. Compared to the previous U1 Lite version, officials specifically highlighted its enhanced control over image-text details.

Practical Applications and Review Feedback

According to hands-on tests and reports from media outlets like APPSO, U1 Pro demonstrated the ability to generate complex visual content such as World Cup match reports and original movie posters. Tech publication Xin Zhiyuan noted that this marks SenseTime's multimodal technology achieving practical potential in complex tasks—transitioning directly from concept comprehension to the delivery of a final visual product.

2026-07-20 ~ 2026-07-21 · 6 related posts

Full story(2 episodes)→