Residual prediction plus a log-determinant regularizer recovers correctly scaled representations

hisspikeness · x · 2026-10-05

The catch: most self-supervised learning pushes representations to spread evenly, destroying exactly the scaling needed for commute-time structure. The authors combine residual prediction with a log-determinant regularizer and prove it recovers the correctly scaled representation.

Related event: CTWM: Commute-Time-Preserving World Model Matches Strong Baseline with Half the Parameters(6 posts)→

Original post →

More from Research

Research channel →