360 Research: Object Segmentation via Video Kinematics

eyishazyer · x · 2026-07-15

This reply highlights a new paper called MoSA from 360 AI Research, which attempts to train models to recognize objects by observing motion in videos, rather than relying on massive amounts of manual annotation.

The post emphasizes the background problem: while models like Meta's SAM are excellent at outlining objects in images, their training costs are exorbitant because they require millions of manually annotated images. MoSA's approach is to use 10,000 hours of unlabeled video to learn "what moves and what is an object," tackling the segmentation/recognition problem at a fraction of the annotation cost.

Related event: MoSA: 360 AI Research learns object segmentation from unlabeled video(5 posts)→

Original post →

More from Research

Research channel →