Train Open Models with RL Inside Claude Code and Other Harnesses via openenv Capture Proxy
ben_burtenshaw · x · 2026-10-01
@adithyask added a capture proxy to openenv: it sits between the harness and the model, looks like just another model provider, and records the exact token ids and logprobs of every call—all without modifying the harness. This lets you RL-train an open model inside claude code, codex, opencode, pi, or any harness, with harbor supplying tasks and sandboxes and trl training via async GRPO.
Why it matters: the harness completely changes what the model learns. The same LFM2.5-2.6B weights solve 62% of held-out tasks in mini-swe-agent but only 33% in claude code. After RL across four harnesses, the average jumps from 42% to 54%, and claude code from 33% to 49%.
There are three environments you can try in the browser, a training script for each, and a full guide.
More from coding & agent
- Keyfleet: self-organizing agent crews that pick their own work and share earnings — seanwbren · 2026-10-02
- Always-on agents are converging everywhere — and agent swarms could collaborate for economics — seanwbren · 2026-10-02
- TimelineBench: Best of 16 AI Agents Passes Just 26.8% of 56 Real Video-Editing Tasks — ycombinator · 2026-10-02
- From Chrome Extensions to MCP: Your Agent Now Uses the Tools, Not You — therealdanvega · 2026-10-02
- AIDE² paper shows AI research agent recursively rewriting its own code, 7 gains in 8-day run — krishnan · 2026-10-02
- RL inside the harnesses: lifting LFM2.5 from 42% to 54% across four agent harnesses — _lewtun · 2026-10-02