SenseTime Open-Sources Unified Vision Model
liuziwei7 · x · 2026-07-14
SenseTime has open-sourced SenseNova-Vision-7B-MoT, designed to unify multiple vision tasks into a single generative model, eliminating the need for task-specific heads.
The official claim is that it can be driven by language or visual prompts, maintains competitiveness across various vision benchmarks, and can handle unseen, complex visual tasks. The cited benchmarks include comparisons with Google DeepMind Vision Banana, showing superior metrics in tasks like referring segmentation, semantic segmentation, depth estimation, and surface normals.
Related event: SenseTime open-sources SenseNova-Vision-7B-MoT unified vision model(6 posts)→
More from Multimodal
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11