Déjà View: Recurrent Transformer for Multi-View 3D Reconstruction
rsasaki0109 · x · 2026-07-20
The open-source project DéjàView (DVLT) has been released. It is a recurrent Transformer model designed for multi-view 3D reconstruction.
- Core Mechanism: Uses recurrent shared frame/global attention combined with discrete depth indices to generate per-pixel rays, depth, confidence, and camera poses from an unstructured set of images.
- Inference Advantage: The model only needs to be trained once, and its refinement steps can be dynamically adjusted during inference. It can match or even surpass larger feedforward baseline models with very few parameters.
- Repository Includes: The DVLT model and four ablation configurations, evaluation wrappers for five baseline models (such as VGGT, Depth-Anything-3, etc.), a training stack based on accelerate + Hydra, as well as Stage-2 fine-tuning schemes and visualization tools.
More from Research
- Chinese AI labs are now treating distillation obfuscation as the top research topic — pmddomingos · 2026-07-22
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22
- enFoldX gains accuracy as AF3 ensemble disagreement rises for non-binders — quaidmorris · 2026-07-22