SenseNova-Vision-7B-MoT Open-Sourced
joemeno · x · 2026-07-15
SenseTime shared an introduction to SenseNova-Vision-7B-MoT: a fully open-source vision model designed to unify multiple vision tasks into a single model. Supported capabilities include:
- Detection, OCR, GUI understanding
- Depth/normal estimation
- Segmentation
- Multi-view tasks
The post also mentions it can define new visual task variants via natural language, recombining capabilities across traditional vision boundaries. Open-source contents include: model weights, the SenseNova-Vision Corpus (a 50M sample subset), and the complete toolchain needed to reproduce the remaining public data. The original post includes links to Hugging Face, GitHub, an online demo, and the technical report.
Related event: SenseTime open-sources SenseNova-Vision-7B-MoT unified vision model(6 posts)→