Tsinghua's CERA-MoA co-evolves routing and LLM agents via RL
Tsinghua · hf · 2026-09-17
Tsinghua researchers introduce CERA-MoA, fixing the disconnect between query routing and agent fine-tuning in Mixture-of-Agents systems:
- Co-evolution: an iterative RL framework lets a dynamic router and independent agent policies evolve together, with routing adapting to agents' changing capabilities during post-training
- Familiarity estimator: mid-layer hidden states assess semantic competence among agents, avoiding full rollout overhead
- Adaptive routing: a cumulative-threshold mechanism activates a tailored minimal agent subset, trading off performance and efficiency
- Targeted training: samples are allocated by evolving competence, promoting specialization
Experiments across domains show CERA-MoA outperforms state-of-the-art static-agent routing and fixed-workflow fine-tuning baselines.
More from coding & agent
- Dev Ships Complete Multiplayer Game Tideball Built Entirely With an LLM — TAbrodi · 2026-09-17
- Multimodal RAG is underused: stop converting audio and video to text first — victorialslocum · 2026-09-17
- Stripe Directory data: merchant playbooks lift agent checkout success from 20/28 to 24/28 — jeff_weinstein · 2026-09-17
- Dev's 3D browser game vibe coding workflow: mockups to WebGPU in a few hours — chongdashu · 2026-09-17
- Mac MCP 2.1.4 ships public endpoint modes, SSRF hardening and transaction undo — bulutarkan · 2026-09-17
- AI filmmaking's hardest problem is no longer video quality — it's continuity — Ok_Low_5536 · 2026-09-17