ICML 论文 ThunderAgent:多轮 Agent RL 吞吐量提升近 2 倍

simran_s_arora · x · 2026-07-30

被 ICML 2026 接收为 Spotlight 的论文 ThunderAgent,针对多轮智能体强化学习中的推理效率瓶颈提出了优化方案。该工作目前已整合至 verl-recipe 的 Dynamo 后端。

核心问题与方案

多轮 Agent 推理常因 KV cache 频繁交换导致 GPU 算力浪费。ThunderAgent 在调度器层面引入了程序感知路由机制来缓解此问题。

性能数据

该系统对提升大规模 Agent 训练与部署的基础设施效率具有显著价值。

所属事件:ThunderAgent系统实现多轮Agent推理近2倍提速(7 条相关)→

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →