PiSSA-style SVD init for LoRA adapters also improves downstream RL training, Trajectory Labs reports
simonguozirui · x · 2026-09-17
The Trajectory Labs team shared a hands-on study: initializing LoRA adapters with the top-k SVD components of pretrained weights (PiSSA, 2024), already known to help SFT, also improves downstream RL training — an idea first suggested in John Schulman's "LoRA Without Regret" blog. They published a blog with clear visuals on why initialization matters and merged the code into SkyRL for others to reproduce. Future directions include LoRA-GA and, for alignment, selecting training directions orthogonal to alignment-critical ones.
More from coding & agent
- Zilliz CTO outlines 'One Data, One Index' architecture to reshape agent retrieval — J_Luan_ · 2026-09-17
- Nelson MCP: open-source extension lets AI agents work inside LibreOffice with 145 tools — quazarzero · 2026-09-17
- Open-source LLM observability tool Opik hits 22k GitHub stars for debugging and evals — dl_weekly · 2026-09-17
- Muse agent books in-terminal hotel and prefills form during Heathrow outage — armand_ruiz · 2026-09-17
- Eval numbers shouldn't be skewed by infra: wall-clock time punishes bad setups — xeophon · 2026-09-17
- Evals shouldn't be skewed by infra: why turn limits break on sub-agents — langstonnashold · 2026-09-17