RelateAnything: a 53M-parameter open-vocabulary CV model that spots object relations in 20ms

blaizedsouza · x · 2026-09-19

RelateAnything, a new open-vocabulary computer vision model, can identify how objects in an image interact — holding, wearing, sitting on — using only raw pixels and bounding boxes. It understands over 19,000 relation words out of the box with zero retraining.

Key details:

Related event: RelateAnything: 53M-parameter model detects object relations in 20ms(2 posts)→

Original post →

More from Embodied

Embodied channel →