SenseTime Open-Sources Unified Vision Model

智东西 · wechat · 2026-07-16

SenseTime has released and fully open-sourced SenseNova-Vision, a unified vision large model designed to consolidate classic vision tasks like detection, segmentation, depth estimation, and 3D reconstruction into a single multimodal generative framework. The article notes it tops the Hugging Face Any-to-Any Leaderboard in overall score, with the model, code, training recipes, and a 50-million-sample corpus all made public.

This approach attempts to end the long-standing "Frankenstein" pattern in vision AI, where different tasks historically relied on separate sets of expert models. SenseNova-Vision unifies the essence of these tasks as generative problems, natively enabling the model to "see" and "understand space." This not only facilitates capability sharing across vision tasks but also allows natural language to directly define new visual tasks.

The article highlights several generalization cases: zero-shot understanding of Minecraft visuals, handling ultra-dense object segmentation, deciphering optical illusions, and seeing through mirror reflections. The author frames this as a step toward general foundation models for vision AI, suggesting it could drive the transition of vision AI from project-based work to platform-level infrastructure.

Related event: SenseTime Open-Sources Unified Vision Model SenseNova-Vision(2 posts)→

Original post →

More from Multimodal

Multimodal channel →