SenseTime Open-Sources Visual Multimodal Model

JaynitMakwana · x · 2026-07-13

SenseTime has released and open-sourced SenseNova-Vision-7B-MoT, promoting the concept of "one model covering major vision tasks."

The content highlights that it reformulates traditional computer vision into a multimodal generative system controllable by language or visual prompts, eliminating the need for task-specific heads.

It also provides benchmark comparisons against Google DeepMind Vision Banana, claiming superior performance on multiple tasks, such as:

The release notes also mention that this is a new generation of visual multimodal model built on SenseTime's decade of experience in vision AI.

Related event: SenseTime open-sources SenseNova-Vision-7B-MoT unified vision model(6 posts)→

Original post →

More from Multimodal

Multimodal channel →