A Deep Derivation of Transformer Init and muP: Why Learning Rates Must Scale With Width

青稞AI · wechat · 2026-09-19

A rigorous long-form article from the Qingke AI community derives the math behind Transformer initialization and extends it to width-scaling learning rates.

Key threads:

Aimed at readers who want to deeply understand training stability and the principles behind muP.

Original post →

More from Research

Research channel →