Harness Learning: Adapting Agent Scaffolds at Test Time Without Touching Model Weights
burny_tech · x · 2026-10-01
Researchers including Ruslan Salakhutdinov and Andrea Zanette (CMU and others) posted "Harness Learning Enables Generalizable Test-Time Adaptation" on arXiv.
- Core idea: train a model to modify an agent's scaffolding program using execution feedback, so the outer harness adapts at test time without updating the underlying model's weights.
- Results: improvements on multi-hop QA and reasoning tasks, enabling generalizable test-time adaptation.
- Why it matters: engineers spend huge amounts of time hand-tuning agent flows, prompts, and tool loops; learned execution feedback offers a blueprint for self-adapting agent frameworks.
- The paper was selected by arXivBangers with an editorial score of 84/100.
Related event: Harness Learning Enables Test-Time Adaptation Without Weight Updates(2 posts)→
More from coding & agent
- Developer uses Claude Opus to turn hundreds of Three.js experiments into a music video — creatoroff · 2026-10-01
- Agent security mindset: least privilege to monitoring, in five layers — goyalshaliniuk · 2026-10-01
- 5 AI agent security risks: browsing, APIs and new attack surfaces — goyalshaliniuk · 2026-10-01
- Dev: I'd rather write dumb utils to parallelize torch trainers than use LLMs for the code — cephaloform · 2026-10-01
- At Jevathon, an AI band where Jev only conducts every 2.7s bar — schwentker · 2026-10-01
- OpenAI's new Decisions API looks a lot like Jev — and 7 hackathon teams built it first — schwentker · 2026-10-01