LoRA weights have been initialized suboptimally: SVD-based init speeds up RL training
burny_tech · x · 2026-09-17
The Trajectory team found that the community's standard LoRA weight initialization is suboptimal: using the SVD of the weight matrix to set both the frozen weights and LoRA initialization yields faster convergence in RL training. The work extends PiSSA and emerged from their exploration of trainable geometries for agentic RL.
Related event: SVD initialization of LoRA accelerates RL training convergence(2 posts)→
More from Research
- Implemented greCAPTCHA Dynamically Generates Multi-Level Questions to Verify Manuscript Authorship — Dr_Atoosa · 2026-09-17
- greCAPTCHA Proposes Measuring Whether Authors Actually Understand Their Own Manuscripts — Dr_Atoosa · 2026-09-17
- greCAPTCHA: testing manuscript-specific understanding to catch AI-written papers — Dr_Atoosa · 2026-09-17
- AI-Generated Submissions Are Flooding Peer Review — A Researcher Proposes Evaluating Process, Not Prose — Dr_Atoosa · 2026-09-17
- "Time Machine Experiments": AI-simulated minds from 1930 as a research method — iyadrahwan · 2026-09-17
- NVIDIA Launches V2D Challenge to Teach Robots From Video — YunzhuLiYZ · 2026-09-17