Stanford's FAMOS infers 3D articulation from sparse partial point clouds feed-forward
Stanford-University · hf · 2026-09-18
Stanford researchers introduce FAMOS, a feed-forward model predicting movable-part segmentation and joint parameters from a sparse, unordered set of partial monocular point clouds, supporting any number of inputs including a single view.
Highlights:
- Multi-state Articulation Transformer with alternating state-wise and global attention aggregates articulation cues across observations.
- An observed articulation span objective supervises each part's motion range across inputs.
- A procedural data generator synthesizes self-annotated assets to overcome dataset scarcity.
Consistent gains over both feed-forward and optimization-based baselines on PartNet-Mobility, ACD, and ArtiCraft-10K.
More from Embodied
- Hugging Face co-founder Thom Wolf shows off iterating robot head prototype — Thom_Wolf · 2026-09-18
- Open-source desk robot Sudo launches: self-hosted AI, $0 subscription, 25 founder units — wateriscoding · 2026-09-18
- LAIA dataset: 15 hours of CARLA driving with human gaze labels for explainable end-to-end AV research — abursuc · 2026-09-18
- Figure runs humanoid robots in 30 rented Bay Area homes with no new training; predicted a $10T company — LinusEkenstam · 2026-09-18
- Vintage robotics video: a handful of PhDs testing modular architectures in the wild — abursuc · 2026-09-18
- Robotics startup founder: hosting events self-selects the hard-to-find ML talent — DominiqueCAPaul · 2026-09-18