Tencent paper: evolving harder environments boosts terminal agents 18 points on Terminal-Bench

rohanpaul_ai · x · 2026-09-09

A new Tencent paper, Environment Evolution for Terminal Agents, argues that as frontier models improve, static synthetic tasks become too easy and stop yielding useful RL signal, while co-evolution methods relying on on-policy rollouts lose effectiveness as the model gets stronger.

The authors propose environment evolution: incrementally increasing environment difficulty off-policy and scheduling evolved environments generation by generation during training to keep learning signals flowing. They derive three difficulty-relevant evolution directions from the multi-turn learning objective and implement them via a multi-agent harness. Rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show consistently harder environments; simple long-horizon RL training on Qwen3.6-27B and Qwen3.6-35B-A3B improves Terminal-Bench 2.1 by 14.4 and 18.0 percentage points respectively.

Related event: Tencent Paper: Evolving Environments Boost Terminal Agents(2 posts)→

Original post →

More from coding & agent

coding & agent channel →