SR-JEPA: Learning Predictive Latent State in 3D Scenes
burny_tech · x · 2026-08-10
SR-JEPA introduces a point-native Joint-Embedding Predictive Architecture (JEPA) for 3D point clouds. Without relying on reconstruction, semantic labels, or language features, the model learns to predict the latent representations of entirely erased entities in a scene using only self-contained 3D EMA targets.
Intervention experiments confirm that the inferred latent state reflects genuine contextual reasoning rather than positional shortcuts. Evaluated on datasets like ARKitScenes, the model significantly outperforms baseline floors in semantic identity imputation and 3D localization, demonstrating a queryable and compositional 3D predictive state.
More from Research
- Indie Dev Builds LLM Consensus Leaderboard Using Esports Ranking Algorithms — 数字生命卡兹克 · 2026-08-10
- ICML Paper: Forcing LLMs to 'Overthink' Leaks Their Hidden Knowledge — PandaAshwinee · 2026-08-10
- Extending Bhattacharyya Coefficients to Power Means for Bayes Error Bounds — FrnkNlsn · 2026-08-10
- GUIDE System Dynamically Generates Multimodal Interactions to Reduce Stress, UIST Paper — _Hao_Zhu · 2026-08-10
- Crime Economists Host Hackathon to Batch Generate Paper Drafts with AI — paulnovosad · 2026-08-10
- PhyLatent: Optimizing JEPA World Model Representations for Better Robot Control — burny_tech · 2026-08-10