Déjà View, a NeurIPS Oral: one looped transformer block matches 3D reconstruction models 8-10x its size
ZGojcic · x · 2026-09-26
The Déjà View paper has been accepted as a NeurIPS Oral. The authors ask whether 3D reconstruction transformers really need a billion parameters, finding most layers do redundant work. Their solution: a single transformer block looped K times. This tiny architecture matches or beats models like VGGT-Ω that are 8-10x larger, with lower compute — evidence of substantial parameter redundancy in 3D reconstruction models.
More from Embodied
- NSF FRR robotics meeting workshop: four talks on skill learning and physical intelligence — YuXiang_IRVL · 2026-09-26
- Prediction: DeepSeek V4 Flash-scale real-time robot vision models coming within weeks — zhaoran_wang · 2026-09-26
- Places Library debuts: 100 high-fidelity real-world 3D environments for embodied AI — yshan2u · 2026-09-26
- microagi x ElevenLabs: Voice is the most human interface for living with robots — animesh_garg · 2026-09-26
- Andy Matuschak builds Printer Friend: voice-to-paper thermal note device — andy_matuschak · 2026-09-26
- Tendon-Driven Robotic Jellyfish Achieves RL-Based Closed-Loop Depth Control — ChongZzZhang · 2026-09-26