ReferTrack tracks language-specified targets with one camera and reaches 89.4% on EVT-Bench

tencent · hf · 2026-07-24

ReferTrack proposes a referring-then-tracking pipeline for embodied visual tracking.

On EVT-Bench, it reports state-of-the-art single-view success rates of 89.4%, 73.3%, and 74.1% on the single-target, distracted, and ambiguity splits.

The authors also say real-world deployments on legged and humanoid robots show strong sim-to-real transfer. Code is available on GitHub.

Original post →

More from Embodied

Embodied channel →