Huawei and CUHK Release Lego-RL for Coding Agent RL Training

青稞AI · wechat · 2026-08-29

Huawei and The Chinese University of Hong Kong (CUHK) released Lego-RL, a reinforcement learning training framework designed for Coding Agents. Its key innovation enables direct RL training on native harnesses (like OpenHands, ClaudeCode) without code modification, addressing the performance degradation seen in traditional methods that alter control flows.

Key Value & Data:

Three Technical Pillars:

Architecture:

Built on verl (training) and Harbor (sandbox execution), connected via AgentLoopWorker and an in-process Proxy. An adapter pattern allows instant support for any harness compatible with OpenAI/Anthropic protocols.

Original post →

More from coding & agent

coding & agent channel →