Salesforce paper: fine-tuning on evolved harnesses makes weak models worse on all 7 tasks
omarsar0 · x · 2026-09-10
A new Salesforce paper co-evolves agent harnesses and models: a harness evolved with a weak model across 7 enterprise tasks, followed by fine-tuning on a stronger expert's trajectories under that harness, dropped performance by 4-30 points on all tasks for Qwen3-Coder and Gemma 4. The same fine-tuning helps under the unevolved harness, pointing to model-harness fit: imitation transfers planning strategies the weak model can't execute. The proposed fix uses a meta-level agent to target failing turns instead of copying whole trajectories.
Related event: Salesforce: Co-evolving Harnesses and Models Shapes Fine-tuning Outcomes(2 posts)→
More from coding & agent
- User burns through 20x Codex quota in two days with heavy subagent use — Aryvyo · 2026-09-10
- Matt Pocock: The Scariest Part of AI Coding Is Not Feeling Your Strategic Mistakes — mattpocockuk · 2026-09-10
- ScreenContext: free open-source Mac recorder that feeds screen videos to AI agents as context — MarcusSchiesser · 2026-09-10
- Dev: Plan 9's 30-year-old 9P protocol fits the external shape of agent systems — remilouf · 2026-09-10
- Personal AGI experiment: a tiny harness, five tools, Markdown, one LLM to give agents memory — dfinke · 2026-09-10
- asc CLI ships per-release changelogs designed for AI agents to consume — rudrank · 2026-09-10