SenseNova-Vision: A 7B Open Model Unifying Segmentation, Depth, and 3D Reconstruction
SandyL925 · reddit · 2026-08-13
SenseNova has open-sourced the 7B vision model SenseNova-Vision (Apache 2.0). Using a Mixture-of-Transformers (MoT) architecture, it treats nearly all computer vision tasks as a single generation problem.
- Core Capabilities: Without task-specific prediction heads, it handles object detection, keypoints, OCR, various segmentations (binary, instance, semantic), and depth/surface normal estimation using a single set of weights.
- Key Highlight: Supports multi-view 3D reconstruction and camera pose estimation. Typically requiring specialized tools like COLMAP, this model accomplishes it with a single prompt.
- Training & Deployment: Trained on 50M instruction-response pairs. A web demo and Hugging Face weights are available. However, hardware requirements are steep: the full web demo needs at least a single 80GB GPU, and benchmarking requires 8x80GB GPUs.
More from Models
- Why Do LLMs Rely on Raw Intelligence Over Effective Communication Post-Training? — matt_slotnick · 2026-08-13
- DeepSeek-V4-Pro-0813 Model Weights Available Again for Download — panchovix · 2026-08-13
- Unsloth Enables Free Fine-Tuning of Muse Glimmer 30B on 24GB VRAM — danielhanchen · 2026-08-13
- Can Open-Source Models Replace Closed Ones? Users Hit by Rate Limits Weigh Switching — SensitiveReading5297 · 2026-08-13
- Researcher Disputes ARC-AGI Uniqueness: Most Benchmarks Show Thinking Model Transitions — scaling01 · 2026-08-13
- AI Assistant Flags Normal Tax Data Update as Fraud in Finance Test — daniel-editide · 2026-08-13