SenseTime Open-Sources Unified Visual Multimodal Model
nikola_mr64990 · x · 2026-07-14
SenseTime has open-sourced SenseNova-Vision-7B-MoT, positioning it as "one model for all major vision tasks." The post highlights that it redefines computer vision as multimodal generation, operable via language or visual prompts without relying on task-specific heads.
Official benchmarks show the model outperforming Google DeepMind's Vision Banana across multiple visual tasks, including:
- Referring expression segmentation: RefCOCOg cIoU 80.3 vs 73.8
- Semantic segmentation: Cityscapes mIoU 71.2 vs 69.9
- Depth estimation: NYUv2 δ1 98.1 vs 94.8
- Normal estimation: NYUv2 mean error 14.4 vs 17.8
The post adds that this direction builds on SenseTime's decade of experience in vision AI, claiming the company has led the Chinese vision AI market for 10 consecutive years.
Related event: SenseTime open-sources SenseNova-Vision-7B-MoT unified vision model(6 posts)→
More from Models
- Grok 4.5 is now free inside Cursor, the popular AI coding IDE — mark_k · 2026-07-21
- GPT often converges on the same near-miss ideas in math problems — yacineMTB · 2026-07-21
- Eno Reyes says model distillation is basically unstoppable — LangChain · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- OpenAI hackathon project stalls as Codex struggles on voice, while Claude spots the issue — ColleenMBrady · 2026-07-21
- Kimi K3 lands exactly on China’s 2-year AI capability trend line — peterwildeford · 2026-07-21