No Input/Output Optimized Params, and No Weight Decay on the LM Head
stochasticchasm · x · 2026-09-11
While examining a model architecture featuring a gigantic engram table, stochasticchasm initially thought it used no Adam optimizer at all, then corrected: some Adam-optimized parameters do remain — but the absence of any input/output optimized parameters is still remarkable. A follow-up in the same thread adds that there is no weight decay on the LM head either.
More from Models
- Multi-agent evals still undecided, but colocated async RL training is catching on — stochasticchasm · 2026-09-11
- Does DeepSeek V4.1-Flash's SWA Bounded Replay sacrifice recall to save KV cache memory? — Top-Handle-5728 · 2026-09-11
- ChatGPT starts inserting ads after each answer, users complain — mansithole6 · 2026-09-11
- Reddit users mourn old coding flow: new models spend 10 minutes overthinking and miss the point — snoosnoosewsew · 2026-09-11
- Forcing models to always max effort is like humans evolving on Adderall, researcher argues — voooooogel · 2026-09-11
- Why Chinese labs distill from Anthropic: Claude's agent data is the scarce training signal — teortaxesTex · 2026-09-11