RelateAnything detects object relations from pixels with a 53M-parameter model at ~20ms per frame

Roger_M_Taylor · x · 2026-09-19

A project called RelateAnything can detect relations between objects in an image — "holding," "behind," "sitting on" and more — using just pixels and bounding boxes.

Key specs: only 53M parameters, 19K+ relation words, no retraining needed, and 20ms per frame. The poster notes spatial relations remain a weak spot but plans to test it with their own detector outputs and video.

Related event: RelateAnything: 53M-parameter model detects object relations in 20ms(2 posts)→

Original post →

More from Research

Research channel →