Salesforce: Imitating Stronger Models' Trajectories Hurts Agent Performance
Salesforce AI's new paper finds that fine-tuning a weak agent on a stronger model's full trajectories degrades performance by 4-30 points, while online error-correction-style fine-tuning is the effective approach.
2026-09-22 ~ 2026-09-22 · 2 related posts
- Salesforce: Fine-tuning a weak model to copy Gemini drops success 15%; correcting its own failures works — rohanpaul_ai · 2026-09-22
- Salesforce: Imitating stronger models drops weak agent performance 4-30 points, on-policy correction wins — rohanpaul_ai · 2026-09-22