Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Qing Zong, Jiayu Liu, Junhao Shen, Zecong Tang, Linsi Wu, Yuxuan Liu, Rui Wang, Zhaowei Wang, Weiqi Wang, Cheng Qian, Xiusi Chen, Yangqiu Song
cs.CL
2026-08-11
This survey organizes co-evolution in agentic systems (multiple components imposing adaptive pressure and reshaping one another) into a progressive three-stage taxonomy: agent-agent (peer adaptation via adversarial, collaborative, or organizational pressure), agent-environment (tasks, feedback, and interaction spaces that change with the agents), and meta co-evolution (the evolution mechanism itself, deciding what, when, how, where, and how to evaluate, becomes evolvable, pointing toward open-endedness). It also covers open challenges in dynamic evaluation, cross-component scaling, and safety governance.
Agentic systems need to keep improving after deployment. Single-entity self-evolution is a main route but is bounded by a static learning context of fixed tasks and feedback. Co-evolution goes further: multiple components impose adaptive pressure on one another and continually reshape each other's subsequent evolution, rather than merely exchanging information or interacting. Despite growing interest, no survey had taken co-evolution as its central axis. This one fills the gap, asking which components co-evolve, how the boundary of evolutionary freedom expands, and what the ultimate form of the process might be.
The authors organize the literature with a progressive three-stage taxonomy whose axis is how the scope of what a system is allowed to evolve expands, progressively shedding human-engineered constraints. Formally, an agentic system S=(A,E) consists of an agent collective A (each agent a made of a model backbone m and a harness h, organized by roles, communication topology, and division of labor) and an environment E; the evolution mechanism driving state transitions is Omega.
Stage 1, agent-agent co-evolution: within a fixed environment, agents are each other's main source of improvement, split into adversarial (red-team attack and defense), collaborative, and organizational adaptation. Stage 2, agent-environment co-evolution: extends adaptation to environmental components, so tasks, feedback, and interaction spaces change with the agents; curriculum learning, agent-task co-generation, and adaptive rewards live here. Stage 3, meta co-evolution: the evolution mechanism itself becomes evolvable. The system decides what, when, how, where, and how to evaluate to evolve, and revises the mechanism through a self-generated process. This recursion offers a path toward open-endedness, where the system keeps producing meaningful novelty with no finite upper bound on adaptive capacity, rather than converging to a fixed endpoint.
For a survey, the results are the landscape and trends it maps. The boundary of evolutionary freedom expands along agents alone, then agents plus environment, then the mechanism itself, matching a gradual withdrawal of human intervention. The lineage runs from biology and games (adversarial self-play) to language agents, curriculum generation, and reward adaptation, with most current systems being local loops (attacker and defender, policy and reward, agent and task generation). The authors are explicit that little work currently meets the strict Stage 3 definition; much of the discussion leans on single-entity meta-evolution as a precursor, showing the mechanism can evolve but not yet coupling that change to a lower-level co-evolving system.
For researchers, this offers a unifying conceptual frame: it realigns work scattered across multi-agent systems, curriculum learning, self-improving agents, and environment generation along a single ruler (the boundary of evolutionary freedom) and draws a clear line between co-evolution and ordinary interaction or information exchange (the requirement of mutually reshaping later evolution). The directions it names are concrete: the future lies not in stronger static agents but in agents that keep improving through co-evolution, which needs matching dynamic evaluation and safety governance. For practitioners tracking where self-improving agents are headed, the three-stage taxonomy is a ready-made map.
The authors flag two main limits. Meta co-evolution is at a very early stage, with limited work meeting its strict definition, so Stage 3 leans heavily on single-entity meta-evolution precursors and is more theoretical than empirical. Safety and governance remain at the level of desiderata with no concrete protocols or mechanisms; they identify failure modes specific to co-evolving systems (evaluator exploitation, partner overfitting, diversity collapse) but develop no safeguards. As a survey it is necessarily bounded by the literature available at the time of writing, and the formal framework is an organizing tool whose predictive power is not validated separately.