CVPR 2026 Best Paper: D4RT enables ultra-fast 4D scene reconstruction
andrew_n_carr · x · 2026-08-30
D4RT, by Google DeepMind et al., won the CVPR 2026 Best Paper. The model uses a unified transformer architecture to jointly infer depth, spatio-temporal correspondence, and full camera parameters from video.
Its core innovation is a novel querying mechanism that sidesteps heavy dense per-frame decoding. The decoding interface allows flexible probing of any 3D point in space and time. This results in highly efficient training and inference, achieving 200+ FPS pose estimation and setting a new state of the art across 4D reconstruction tasks.
More from Multimodal
- H3 REF2VID Demo: Video Generation Without Masking — AthleteEducational63 · 2026-08-30
- GPT-6 generates stunning Pagoda Voxel art — ChrisGPT · 2026-08-30
- Seedance 2.5 outperforms Grok 1.5 in art direction test — ChrisGPT · 2026-08-30
- Dev builds custom ComfyUI node for continuous, length-unlimited MiniMax H3 video generation — rynaleopard · 2026-08-30
- HR Endless Sampler: open-source ComfyUI node renders any-length H3 videos on 16GB VRAM — rhradec · 2026-08-30
- MiniMax-H3-Longvideos hits Hugging Face trending: long-form text-to-video with synced audio — Smite79 · 2026-08-30