SR-JEPA: Learning Predictive Latent State in 3D Scenes

burny_tech · x · 2026-08-10

SR-JEPA introduces a point-native Joint-Embedding Predictive Architecture (JEPA) for 3D point clouds. Without relying on reconstruction, semantic labels, or language features, the model learns to predict the latent representations of entirely erased entities in a scene using only self-contained 3D EMA targets.

Intervention experiments confirm that the inferred latent state reflects genuine contextual reasoning rather than positional shortcuts. Evaluated on datasets like ARKitScenes, the model significantly outperforms baseline floors in semantic identity imputation and 3D localization, demonstrating a queryable and compositional 3D predictive state.

Original post →

More from Research

Research channel →