No Input/Output Optimized Params, and No Weight Decay on the LM Head

stochasticchasm · x · 2026-09-11

While examining a model architecture featuring a gigantic engram table, stochasticchasm initially thought it used no Adam optimizer at all, then corrected: some Adam-optimized parameters do remain — but the absence of any input/output optimized parameters is still remarkable. A follow-up in the same thread adds that there is no weight decay on the LM head either.

Related event: Muon becomes the default optimizer as new model drops MTP and applies QAT to KV cache(7 posts)→

Original post →

More from Models

Models channel →