Tencent paper: environment evolution generates harder RL environments without watching the agent
omarsar0 · x · 2026-09-05
A Tencent paper tackles environment supply as the main bottleneck for agent RL. Prior methods derive environments from weaknesses in the agent's own rollouts, inheriting blind spots and weakening as the agent improves. Environment evolution instead derives three difficulty-raising transformations directly from the multi-turn training objective and applies them generation by generation on a fixed schedule. Hy4 preview, Claude Opus 5 and GPT-5.6 Sol all score worse on evolved environments, validating the approach.
More from Research
- Tivadar Danka Maps the Knowledge Graph of Machine Learning, From Math Foundations to SOTA — TivadarDanka · 2026-09-05
- Researcher Maps the Machine Learning Knowledge Graph, Foundations to Frontier — TivadarDanka · 2026-09-05
- Thinking Machines open-sources RL recipe for fine-tuning models as event probability forecasters — clarejtbirch · 2026-09-05
- Anthropic posts a complete Lean 4 machine-checked proof of Fermat's Last Theorem — scaling01 · 2026-09-05
- Failure modes found only by running coding agents unattended for months — Fragrant_Yoghurt1135 · 2026-09-05
- 2D-RoPE for image models is easier in the complex representation, says arohan — _arohan_ · 2026-09-05