SenseTime open-sources SenseNova-Vision, a unified model for detection to 3D reconstruction
socialwithaayan · x · 2026-07-29
SenseTime open-sourced SenseNova-Vision, a unified multimodal model that tries to collapse a long list of computer-vision tasks into one set of weights.
- The repo presents the model as “Vision as Unified Multimodal Generation.”
- It covers detection, segmentation, depth, surface normals, OCR, keypoints, and 3D reconstruction.
- The author claims a single text interface can drive all these tasks, instead of stitching together multiple specialist models.
- The post also notes real-world limits: on a living-room scene, the model produced mostly clean masks and reasonable reasoning, but still mislabeled a couch cushion as a book.
More from Multimodal
- A 3D pose editor turns skeletons into depth maps for Krea 2 and Z-Image — ashishsanu · 2026-07-29
- ComfyUI node merges two Krea 2 bf16 models and quantizes them to INT8 or INT4 — Away_Exam_4586 · 2026-07-29
- Flux3 shows stronger style preservation for animated children's book visuals — arnicas · 2026-07-29
- Professor Tests Flux 3: Extremely High Prompt Adherence — emollick · 2026-07-29
- AI image generators can infer a body pose from one photo, and users find it creepy — Numerous_Temporary11 · 2026-07-29
- Flux3 follows Gertrude Abercrombie style and owl prompts in a new test — arnicas · 2026-07-29