SeeSE3 finds 3D structure emerging in frozen vision features and camera-pose alignment

ducha_aiki · x · 2026-07-21

The post points to SeeSE3: Emergence of 3D Space in Vision Features, a paper looking at whether frozen vision encoders contain recoverable 3D structure.

Key takeaways from the thread and figures:

The analysis section says the answer to the original question is effectively yes: a motionless observer can discover space, but only up to a nonlinear unwrapping that a simple adapter can undo. It also emphasizes that probe design matters: Lie-algebra linearization and a Siamese structure outperform alternatives, and global decodability depends on the curvature of the feature manifold.

Original post →

More from Research

Research channel →