Tencent paper: continuously harder task environments beat co-evolution, +8.6pp on Terminal-Bench
rohanpaul_ai · x · 2026-09-09
A new Tencent paper argues that as agents improve, training tasks cannot stay still: continuously harder environments outperform co-evolution.
- Key insight: synthetic terminal tasks eventually become too easy; once solved reliably, they stop providing useful RL signal.
- Method: instead of building tasks around model failures, evolve the tasks themselves—less familiar setups, rarer required skills, more steps—verifying each harder version before introducing it.
- Results: on Terminal-Bench 2.1, Qwen3.6-27B hit 71.5% vs 62.9% with co-evolution; Qwen3.6-35B-A3B hit 64.9% vs 55.1%. Tested with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol.
Related event: Tencent Paper: Evolving Environments Boost Terminal Agents(2 posts)→
More from coding & agent
- Stop using random multi-agent topologies: how to audit wasted subagent delegation — keyanzhang · 2026-09-09
- opencode v1.18.30 Adds Astra System Prompt for GPT-6, Fixes Bedrock DeepSeek IDs — opencode-agent[bot] · 2026-09-09
- Abacus.AI teases near-free LLM targeting long-running personal agentic loops, launching Thursday — bindureddy · 2026-09-09
- huashu-chrome: Open-Source MCP + Chrome Extension Lets Agents Drive Your Logged-in Browser — AlchainHust · 2026-09-09
- GPT-6 Astra instantly finished browser agent tasks Claude Code couldn't, user reports — AlchainHust · 2026-09-09
- How Do You Verify an Agent Fix When It Only Fails One Run in Twenty? — Such-Process5697 · 2026-09-09