New Research Tackles LoRA's Random Down-Projection Bottleneck for Faster Fine-Tuning

burkov · x · 2026-09-07

Andriy Burkov summarizes new research on LoRA's optimization flaws: because the up-projection is zero-initialized, early training relies entirely on the randomly initialized down-projection, which creates unbalanced implicit learning rates across input dimensions and weak initial gradients, slowing convergence. The authors propose regularizing the down-projection to fix these distortions while keeping LoRA's parameter-efficiency benefits.

Related event: New Research Explains LoRA's Slow Convergence, Proposes One-Line NoRA Fix(3 posts)→

Original post →

More from Research

Research channel →