SenseTime Open-Sources Unified Vision Model

liuziwei7 · x · 2026-07-14

SenseTime has open-sourced SenseNova-Vision-7B-MoT, designed to unify multiple vision tasks into a single generative model, eliminating the need for task-specific heads.

The official claim is that it can be driven by language or visual prompts, maintains competitiveness across various vision benchmarks, and can handle unseen, complex visual tasks. The cited benchmarks include comparisons with Google DeepMind Vision Banana, showing superior metrics in tasks like referring segmentation, semantic segmentation, depth estimation, and surface normals.

Related event: SenseTime open-sources SenseNova-Vision-7B-MoT unified vision model(6 posts)→

Original post →

More from Multimodal

Multimodal channel →