Tencent paper: environment evolution generates harder RL environments without watching the agent

omarsar0 · x · 2026-09-05

A Tencent paper tackles environment supply as the main bottleneck for agent RL. Prior methods derive environments from weaknesses in the agent's own rollouts, inheriting blind spots and weakening as the agent improves. Environment evolution instead derives three difficulty-raising transformations directly from the multi-turn training objective and applies them generation by generation on a fixed schedule. Hy4 preview, Claude Opus 5 and GPT-5.6 Sol all score worse on evolved environments, validating the approach.

Original post →

More from Research

Research channel →