SeeSE3 finds 3D structure emerging in frozen vision features and camera-pose alignment
ducha_aiki · x · 2026-07-21
The post points to SeeSE3: Emergence of 3D Space in Vision Features, a paper looking at whether frozen vision encoders contain recoverable 3D structure.
Key takeaways from the thread and figures:
- Depth-related and DINO-family features show camera-pose alignment, with stronger alignment after head fine-tuning.
- The authors compare short- and long-range spatial topology on ScanNet and find that explicit geometric models like DUST3R and MoGe score highly.
- Among self-supervised models, DINO-family encoders exhibit emerging spatial topology, while video models such as 4DS and RVM struggle to keep a consistent spatial embedding.
- A Poincaré adapter is evaluated across different frame strides; the paper argues that changes in latent features can be linearly mapped to camera pose changes.
The analysis section says the answer to the original question is effectively yes: a motionless observer can discover space, but only up to a nonlinear unwrapping that a simple adapter can undo. It also emphasizes that probe design matters: Lie-algebra linearization and a Siamese structure outperform alternatives, and global decodability depends on the curvature of the feature manifold.
More from Research
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- BlackboxNLP 2026 is recruiting extra reviewers after a high submission volume — hanjie_chen · 2026-07-22
- AWS shows self-distilled reasoning can preserve math and coding skills during SFT — AWS ML Blog · 2026-07-22
- UI2App shows screenshot fidelity still lags real interaction recovery — Grace Man Chen · 2026-07-22