Déjà View, a NeurIPS Oral: one looped transformer block matches 3D reconstruction models 8-10x its size

ZGojcic · x · 2026-09-26

The Déjà View paper has been accepted as a NeurIPS Oral. The authors ask whether 3D reconstruction transformers really need a billion parameters, finding most layers do redundant work. Their solution: a single transformer block looped K times. This tiny architecture matches or beats models like VGGT-Ω that are 8-10x larger, with lower compute — evidence of substantial parameter redundancy in 3D reconstruction models.

Original post →

More from Embodied

Embodied channel →