TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
rsasaki0109 · x · 2026-08-15
Researchers from KAIST and Google DeepMind propose TrackCraft3R, the first method to repurpose a video diffusion transformer (video DiT) as a feed-forward dense 3D tracker. Given a monocular video and its frame-anchored reconstruction pointmap, TrackCraft3R predicts a reference-anchored tracking pointmap that follows every pixel of the first frame across time in a single forward pass, along with visibility. Through designs like dual-latent representation, it addresses the mismatch between frame-anchored formulation of video DiTs and reference-anchored tracking. Experiments show state-of-the-art dense 3D tracking performance on real-world videos.
More from Research
- Insilico Medicine unveils Virtual Aging Cell platform with age as a core variable — MaxUnfried · 2026-08-15
- "Mathematical and Statistical Foundations of AI" Book Released with AI Tutor — burkov · 2026-08-15
- NSF awards $20M to Northwestern for first public AI protein-engineering cloud lab — MichaelCJewett · 2026-08-15
- GPU program search is hindered by massive search space, requires abstraction — spikedoanz · 2026-08-15
- Why DeepSeek's dsh/cordis is a big deal: LH tasks and Harness meta tuning — EstablishmentOdd785 · 2026-08-15
- Stanford Virtual Embryo Challenge draws 401 researchers and 295 teams in first week — anshulkundaje · 2026-08-15