TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

rsasaki0109 · x · 2026-08-15

Researchers from KAIST and Google DeepMind propose TrackCraft3R, the first method to repurpose a video diffusion transformer (video DiT) as a feed-forward dense 3D tracker. Given a monocular video and its frame-anchored reconstruction pointmap, TrackCraft3R predicts a reference-anchored tracking pointmap that follows every pixel of the first frame across time in a single forward pass, along with visibility. Through designs like dual-latent representation, it addresses the mismatch between frame-anchored formulation of video DiTs and reference-anchored tracking. Experiments show state-of-the-art dense 3D tracking performance on real-world videos.

Original post →

More from Research

Research channel →