ReferTrack tracks language-specified targets with one camera and reaches 89.4% on EVT-Bench
tencent · hf · 2026-07-24
ReferTrack proposes a referring-then-tracking pipeline for embodied visual tracking.
- It tackles EVT with a single forward-facing camera.
- The model first selects the target from indexed bounding boxes, then predicts tracking waypoints.
- To preserve motion cues, it keeps a sliding window of previously selected boxes and injects their geometry through TVBI tokens.
- It also co-trains on a custom Refer-QA dataset to improve target identification.
On EVT-Bench, it reports state-of-the-art single-view success rates of 89.4%, 73.3%, and 74.1% on the single-target, distracted, and ambiguity splits.
The authors also say real-world deployments on legged and humanoid robots show strong sim-to-real transfer. Code is available on GitHub.
More from Embodied
- IGGT4D brings streaming 4D instance-grounded geometry to dynamic scene understanding — zhenjun_zhao · 2026-07-24
- GLAM-SLAM pairs ORB-SLAM2 with Gaussian mapping for real-time large-scale reconstruction — zhenjun_zhao · 2026-07-24
- MIT report: nuclear plants are shifting toward supervised autonomous control — nordicinst · 2026-07-24
- Black Forest Labs video models are being used to automate Audi manufacturing — hsu_byron · 2026-07-24
- Hong Kong prepares for its first test run of fully driverless vehicles — Baidu_Inc · 2026-07-24
- Neuralink says trial participants with paralysis can drive wheelchairs by thought — Polymarket · 2026-07-24