LoRA weights have been initialized suboptimally: SVD-based init speeds up RL training

burny_tech · x · 2026-09-17

The Trajectory team found that the community's standard LoRA weight initialization is suboptimal: using the SVD of the weight matrix to set both the frozen weights and LoRA initialization yields faster convergence in RL training. The work extends PiSSA and emerged from their exploration of trainable geometries for agentic RL.

Related event: SVD initialization of LoRA accelerates RL training convergence(2 posts)→

Original post →

More from Research

Research channel →