Residual prediction plus a log-determinant regularizer recovers correctly scaled representations
hisspikeness · x · 2026-10-05
The catch: most self-supervised learning pushes representations to spread evenly, destroying exactly the scaling needed for commute-time structure. The authors combine residual prediction with a log-determinant regularizer and prove it recovers the correctly scaled representation.
More from Research
- Dream4ACT Unifies Video-Action Modeling Across Robot Embodiments, Hitting 89% on RoboTwin 2.0 — Xiangyu Zhu · 2026-10-05
- SMI Brings Understanding-Driven Spatial Memory Management to Long-Video World Models — Ying Yang · 2026-10-05
- LVMT Sets New SOTA in Long-Term Video Segmentation While Running 10X Faster — tue-mps · 2026-10-05
- Schmidhuber: Google's 2017 Transformer Builds on His 1991 Linear Attention Work — SchmidhuberAI · 2026-10-05
- Near-identical image scores, huge gaps: AI denoising must serve science, not looks — bravo_abad · 2026-10-05
- SDECast: Neural SDEs Bring Continuous-Time Probabilistic Weather Forecasts Out to 5 Days — canaesseth · 2026-10-05