SenseTime SenseNova-Vision: Unifies Visual Tasks via Multimodal Generation
rsasaki0109 · x · 2026-08-24
SenseTime's SenseNova-Vision formulates computer vision as a unified multimodal generation task, handling heterogeneous visual tasks through native text and image generation spaces. Natural language and visual prompts specify tasks, targets, and output schemas.
- Text Generation: Outputs symbolic records like categories, boxes, OCR strings, and camera parameters.
- Image Generation: Handles dense spatial targets such as segmentation masks, depth maps, and point maps.
- Mixed Responses: Supports compositional tasks combining symbolic and dense outputs.
This shared formulation allows a single model to handle diverse visual tasks without task-specific architectures.
More from Multimodal
- UniSpace: Unified Visual Representation without VAE — meituan-longcat · 2026-08-24
- Wan 3.0 demo impresses with lifelike faces and stable motion — JaynitMakwana · 2026-08-24
- Grafting Krea2 into MiniMax H3: improving texture via component transplant — Key-Philosopher-9327 · 2026-08-24
- Cinematic FLUX 3 Test: Motorcyclist in the Desert with Full Prompt — LudovicCreator · 2026-08-24
- Struggling with hand and finger generation in Minimax H3? — Glittering-Cold-2981 · 2026-08-24
- Conflict Between Minimax H3 Motion Context and Latent Upscaling — Drock_belg · 2026-08-24