Asymmetric Optimization Distances in Weight Scales

torchcompiled · x · 2026-07-15

The author argues that many components in neural networks are inherently multiplicative.

Using LayerNorm scaling as an example:

However, from an optimization perspective, the "optimization distance" between 0.5 and 1.0 is not symmetric to that between 1.0 and 2.0, which is why he considered this new weight parameterization.

Related event: SEL Weight Reparameterization: ~1.42x Training Speedup, Folds Back to Standard Weights(7 posts)→

Original post →

More from Research

Research channel →