SenseTime Open-Sources Unified Vision Model
liuziwei7 · x · 2026-07-14
SenseTime has open-sourced SenseNova-Vision-7B-MoT, designed to unify multiple vision tasks into a single generative model, eliminating the need for task-specific heads.
The official claim is that it can be driven by language or visual prompts, maintains competitiveness across various vision benchmarks, and can handle unseen, complex visual tasks. The cited benchmarks include comparisons with Google DeepMind Vision Banana, showing superior metrics in tasks like referring segmentation, semantic segmentation, depth estimation, and surface normals.
Related event: SenseTime open-sources SenseNova-Vision-7B-MoT unified vision model(6 posts)→
More from Multimodal
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22