SenseTime Open-Sources Unified Visual Multimodal Model

nikola_mr64990 · x · 2026-07-14

SenseTime has open-sourced SenseNova-Vision-7B-MoT, positioning it as "one model for all major vision tasks." The post highlights that it redefines computer vision as multimodal generation, operable via language or visual prompts without relying on task-specific heads.

Official benchmarks show the model outperforming Google DeepMind's Vision Banana across multiple visual tasks, including:

The post adds that this direction builds on SenseTime's decade of experience in vision AI, claiming the company has led the Chinese vision AI market for 10 consecutive years.

Related event: SenseTime open-sources SenseNova-Vision-7B-MoT unified vision model(6 posts)→

Original post →

More from Models

Models channel →