HuggingFace tackles harness overfitting with multi-harness RL across Claude Code, Codex and more

huggingface · x · 2026-10-05

HuggingFace researchers propose multi-harness RL to fix a known gap: an open-weight model trained inside one agent harness often loses accuracy or emits invalid tool calls when moved to another, because it learns that harness's tool names, formats and control flow rather than the task. They trained through four harnesses simultaneously — Claude Code, Codex, OpenCode and Mini-SWE-Agent — connected via three open systems: OpenEnv provides a standard interface between harnesses, RL environments and trainers, with its capture proxy in the middle. Also linked: the author's deep-dive article on how Anthropic, OpenAI, Perplexity and LangChain build agent harnesses.

Related event: HuggingFace Turns Claude Code, Codex and 8 Other Coding Tools into RL Training Environments(2 posts)→

Original post →

More from coding & agent

coding & agent channel →