CMU's TrackEverything tracks all visible points in 1000+ frame videos within 40GB GPU memory

CarnegieMellonU · hf · 2026-09-28

CMU researchers introduce TrackEverything, a 3D point tracker that breaks the trade-off between tracking sparse points over long horizons and tracking all points only in short clips, by representing videos as persistent 3D scene tracks in world coordinates — decoupling model complexity from video duration.

It is reportedly the first 3D tracker to track all visible points in videos exceeding 1000 frames within 40GB GPU memory, outperforming open-source dense 3D trackers by over 20% APD on TAPVid-3D short clips while staying competitive with sparse trackers on long sequences.

Original post →

More from Research

Research channel →