RelateAnything: a 53M-parameter open-vocabulary CV model that spots object relations in 20ms
blaizedsouza · x · 2026-09-19
RelateAnything, a new open-vocabulary computer vision model, can identify how objects in an image interact — holding, wearing, sitting on — using only raw pixels and bounding boxes. It understands over 19,000 relation words out of the box with zero retraining.
Key details:
- Just 53 million parameters, roughly 20ms per frame, viable for real-time video streaming and robotics
- Relies purely on visual context and geometry rather than guessing from object names
- Weak spot: spatial relations remain unreliable
Related event: RelateAnything: 53M-parameter model detects object relations in 20ms(2 posts)→
More from Embodied
- Fruit fly brain connectome drives an eBay Vector robot with 166,700 simulated neurons — sull · 2026-09-19
- Over 100M Americans Wear Sensors — But Do HRV and Readiness Scores Hold Up? — EricTopol · 2026-09-19
- Indie robot U-BOT day 32: prototype control board mount ready for motor torque testing — _Stocko_ · 2026-09-19
- Tesla AI5 chip enters trial production on Samsung's 2nm Texas fab, mass output by 2027 — XFreeze · 2026-09-19
- Real world isn't a simulator: why autonomous AI struggles to go physical — AlexTensor · 2026-09-19
- Robotics RL is brutally hard: one practitioner's list of a dozen failure modes — Scobleizer · 2026-09-19