Microsoft’s OpenForgeRL trains harness-native agents end to end in any environment
Fighterdan · hf · 2026-07-25
Microsoft researchers introduce OpenForgeRL, an open-source framework for training harness-native agents end to end in any environment.
What it does
- Adds a lightweight proxy to record harness model calls as RL training data.
- Uses a Kubernetes orchestrator so each rollout runs in its own remote container.
- Lets researchers train agents inside real harnesses such as Claude Code, Codex, OpenClaw, and GUI/browser/computer-use setups.
Reported results
- OpenForgeClaw: 31.7 pass^3 on ClawEval and 55.9 pass@3 on ClawEval, plus 33.7 on QwenClawBench.
- OpenForgeGUI: 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager.
- The paper says these models beat similarly sized open baselines on nearly all benchmarks, and in GUI tasks can match or exceed models several times larger.
Takeaways
- Harness choice materially changes agent behavior.
- RL improves reliability, self-verification, tool coverage, and multi-step plan completion.
- Error recovery is still weak.
Related event: Microsoft Open-Sources OpenForgeRL for End-to-End Agent Training(4 posts)→
More from coding & agent
- The browser main thread is expensive: a practical guide to JavaScript and CSS animation cost — jh3yy · 2026-09-11
- Inspired by OpenAI's 10,000-agent run, dev open-sources a crowdsourced agent problem-solving platform — Benjaminsen · 2026-09-11
- Lucid: open-source Mac app keeps your laptop awake only while AI agents run — Pitiful_Hedgehog_600 · 2026-09-11
- banteg's snail project crowdsources AI agents to finish matching Snail Mail's 20 remaining functions — banteg · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11