Tencent paper: evolving harder environments boosts terminal agents 18 points on Terminal-Bench
rohanpaul_ai · x · 2026-09-09
A new Tencent paper, Environment Evolution for Terminal Agents, argues that as frontier models improve, static synthetic tasks become too easy and stop yielding useful RL signal, while co-evolution methods relying on on-policy rollouts lose effectiveness as the model gets stronger.
The authors propose environment evolution: incrementally increasing environment difficulty off-policy and scheduling evolved environments generation by generation during training to keep learning signals flowing. They derive three difficulty-relevant evolution directions from the multi-turn learning objective and implement them via a multi-agent harness. Rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show consistently harder environments; simple long-horizon RL training on Qwen3.6-27B and Qwen3.6-35B-A3B improves Terminal-Bench 2.1 by 14.4 and 18.0 percentage points respectively.
Related event: Tencent Paper: Evolving Environments Boost Terminal Agents(2 posts)→
More from coding & agent
- Stop using random multi-agent topologies: how to audit wasted subagent delegation — keyanzhang · 2026-09-09
- opencode v1.18.30 Adds Astra System Prompt for GPT-6, Fixes Bedrock DeepSeek IDs — opencode-agent[bot] · 2026-09-09
- Abacus.AI teases near-free LLM targeting long-running personal agentic loops, launching Thursday — bindureddy · 2026-09-09
- huashu-chrome: Open-Source MCP + Chrome Extension Lets Agents Drive Your Logged-in Browser — AlchainHust · 2026-09-09
- GPT-6 Astra instantly finished browser agent tasks Claude Code couldn't, user reports — AlchainHust · 2026-09-09
- How Do You Verify an Agent Fix When It Only Fails One Run in Twenty? — Such-Process5697 · 2026-09-09