SenseTime Open-Sources Unified Visual Multimodal Model
nikola_mr64990 · x · 2026-07-14
SenseTime has open-sourced SenseNova-Vision-7B-MoT, positioning it as "one model for all major vision tasks." The post highlights that it redefines computer vision as multimodal generation, operable via language or visual prompts without relying on task-specific heads.
Official benchmarks show the model outperforming Google DeepMind's Vision Banana across multiple visual tasks, including:
- Referring expression segmentation: RefCOCOg cIoU 80.3 vs 73.8
- Semantic segmentation: Cityscapes mIoU 71.2 vs 69.9
- Depth estimation: NYUv2 δ1 98.1 vs 94.8
- Normal estimation: NYUv2 mean error 14.4 vs 17.8
The post adds that this direction builds on SenseTime's decade of experience in vision AI, claiming the company has led the Chinese vision AI market for 10 consecutive years.
Related event: SenseTime open-sources SenseNova-Vision-7B-MoT unified vision model(6 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11