Salesforce paper: fine-tuning on evolved harnesses makes weak models worse on all 7 tasks

omarsar0 · x · 2026-09-10

A new Salesforce paper co-evolves agent harnesses and models: a harness evolved with a weak model across 7 enterprise tasks, followed by fine-tuning on a stronger expert's trajectories under that harness, dropped performance by 4-30 points on all tasks for Qwen3-Coder and Gemma 4. The same fine-tuning helps under the unevolved harness, pointing to model-harness fit: imitation transfers planning strategies the weak model can't execute. The proposed fix uses a meta-level agent to target failing turns instead of copying whole trajectories.

Related event: Salesforce: Co-evolving Harnesses and Models Shapes Fine-tuning Outcomes(2 posts)→

Original post →

More from coding & agent

coding & agent channel →