kalomaze: In LoRA, Weight Decay Behaves Like Regularizing Toward the Base Model
kalomaze · x · 2026-10-04
kalomaze points out a subtle insight: in the LoRA setting, weight decay's semantic effect is much closer to "regularizing toward the base model's original weights" than "penalizing absolute weight magnitudes." Because LoRA trains only low-rank deltas, the regularization effectively pulls updates toward the base weights rather than toward zero — a useful reframing for understanding regularization in fine-tuning.
More from Research
- Mike Levin to discuss causal architecture dynamics in pre-replicator catalytic networks — drmichaellevin · 2026-10-04
- Schmidhuber's Formal Theory of Fun and Creativity Dates Back to a 2009 Japanese Journal Paper — hardmaru · 2026-10-04
- Fighting AI slop: publishers and researchers push back on machine-written science papers — Symbiot10000 · 2026-10-04
- Equinor finds 27 oil discoveries by running old seismic data through new AI models — YvesMulkers · 2026-10-04
- Guest post on Terence Tao's blog: how AI does, and does not, change how I do math — stevenstrogatz · 2026-10-04
- Book on Lean proof assistant hits #2 on WSJ's weekly reading list — stevenstrogatz · 2026-10-04