CMU's TrackEverything tracks all visible points in 1000+ frame videos within 40GB GPU memory
CarnegieMellonU · hf · 2026-09-28
CMU researchers introduce TrackEverything, a 3D point tracker that breaks the trade-off between tracking sparse points over long horizons and tracking all points only in short clips, by representing videos as persistent 3D scene tracks in world coordinates — decoupling model complexity from video duration.
- Voxel dedup: merges co-located tracks at sliding-window boundaries to prevent redundant accumulation.
- Decomposed tracking: an endpoint refiner predicts each point's destination and static/dynamic class, then a lightweight trajectory refiner decodes dense trajectories only for dynamic points.
- 3D WAFT: replaces memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud.
It is reportedly the first 3D tracker to track all visible points in videos exceeding 1000 frames within 40GB GPU memory, outperforming open-source dense 3D trackers by over 20% APD on TAPVid-3D short clips while staying competitive with sparse trackers on long sequences.
More from Research
- Yoav Goldberg: 29 arXiv papers on JEV just two weeks after its limited release is 'not healthy' — yoavgo · 2026-09-28
- Why Data Efficiency Is Humans' Last Edge — Yet Massive Data Feeding Stays Cheapest — himanshustwts · 2026-09-28
- Harness-Zero: PKU, Google and HKUST Distill Agent Harnesses into Model Weights — AxSaucedo · 2026-09-28
- MultiFlow: coupled flow matching predicts single-cell perturbation responses in unseen cellular contexts — burny_tech · 2026-09-28
- Self-Replicating Programs Emerge from Random Noise in Open-Source Browser Remake of BRLabs' Computational Life — zzznah · 2026-09-28
- Pop-Corn preprint targets unseen-perturbation cell composition shifts in Perturb-seq — burny_tech · 2026-09-28