MoSA: Learning Visual Concepts From Video Motion
ZabihullahAtal · x · 2026-07-15
This post introduces MoSA from the 360 AI Research Institute. Instead of relying on massive manual image annotations, it attempts to learn object concepts from motion in unlabeled videos.
Key details include:
- Utilizes roughly 10,000 hours of unlabeled video.
- Generates over 21 million pseudo-labels.
- Aims to reduce reliance on manual annotation and explore a more scalable training path for foundation vision models.
- The authors claim it outperforms UnSAM and is comparable to SAM.
The post also notes that this reflects a broader trend: Chinese AI companies are increasing their presence at top-tier conferences like ECCV.
Related event: MoSA: 360 AI Research learns object segmentation from unlabeled video(5 posts)→
More from Research
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- BlackboxNLP 2026 is recruiting extra reviewers after a high submission volume — hanjie_chen · 2026-07-22
- AWS shows self-distilled reasoning can preserve math and coding skills during SFT — AWS ML Blog · 2026-07-22
- UI2App shows screenshot fidelity still lags real interaction recovery — Grace Man Chen · 2026-07-22