ReferTrack tracks language-specified targets with one camera and reaches 89.4% on EVT-Bench
tencent · hf · 2026-07-24
ReferTrack proposes a referring-then-tracking pipeline for embodied visual tracking.
- It tackles EVT with a single forward-facing camera.
- The model first selects the target from indexed bounding boxes, then predicts tracking waypoints.
- To preserve motion cues, it keeps a sliding window of previously selected boxes and injects their geometry through TVBI tokens.
- It also co-trains on a custom Refer-QA dataset to improve target identification.
On EVT-Bench, it reports state-of-the-art single-view success rates of 89.4%, 73.3%, and 74.1% on the single-target, distracted, and ambiguity splits.
The authors also say real-world deployments on legged and humanoid robots show strong sim-to-real transfer. Code is available on GitHub.
More from Embodied
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11