Pretrained Weights Exhibit Heavy-Tailed Distribution

torchcompiled · x · 2026-07-15

The author states that this exploration stems from a long-standing puzzle: while preconditioning helps, it doesn't truly account for the "optimization distance" certain parameters need to cover.

He also mentions observing a heavy-tailed weight distribution in pretrained models, which might be related to this parameter scaling and optimization difficulty.

Related event: SEL Weight Reparameterization: ~1.42x Training Speedup, Folds Back to Standard Weights(7 posts)→

Original post →

More from Research

Research channel →