SenseNova-Vision-7B-MoT Released
liuziwei7 · x · 2026-07-10
SenseNova-Vision-7B-MoT is now available on ModelScope as a unified multimodal generative model designed for computer vision tasks.
According to the post, it outperforms general vision models across tasks like object detection, semantic segmentation, visual grounding, and depth estimation. It outputs results via text, images, or a mix of both, and was trained on SenseNova-Vision-Corpus-50M under the CC BY-NC 4.0 license.
Related event: SenseNova-Vision Unifies Vision Tasks With 7B-MoT(4 posts)→
More from Multimodal
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22