CVPR 2026 Best Paper: D4RT enables ultra-fast 4D scene reconstruction

andrew_n_carr · x · 2026-08-30

D4RT, by Google DeepMind et al., won the CVPR 2026 Best Paper. The model uses a unified transformer architecture to jointly infer depth, spatio-temporal correspondence, and full camera parameters from video.

Its core innovation is a novel querying mechanism that sidesteps heavy dense per-frame decoding. The decoding interface allows flexible probing of any 3D point in space and time. This results in highly efficient training and inference, achieving 200+ FPS pose estimation and setting a new state of the art across 4D reconstruction tasks.

Original post →

More from Multimodal

Multimodal channel →