Train Open Models with RL Inside Claude Code and Other Harnesses via openenv Capture Proxy

ben_burtenshaw · x · 2026-10-01

@adithyask added a capture proxy to openenv: it sits between the harness and the model, looks like just another model provider, and records the exact token ids and logprobs of every call—all without modifying the harness. This lets you RL-train an open model inside claude code, codex, opencode, pi, or any harness, with harbor supplying tasks and sandboxes and trl training via async GRPO.

Why it matters: the harness completely changes what the model learns. The same LFM2.5-2.6B weights solve 62% of held-out tasks in mini-swe-agent but only 33% in claude code. After RL across four harnesses, the average jumps from 42% to 54%, and claude code from 33% to 49%.

There are three environments you can try in the browser, a training script for each, and a full guide.

Original post →

More from coding & agent

coding & agent channel →