ThunderAgent (ICML Spotlight): Nearly 2x Throughput for Multi-Turn Agent RL
simran_s_arora · x · 2026-07-30
Proposed by Together AI, ThunderAgent tackles inference scheduling bottlenecks in multi-turn Agent reinforcement learning. It has been accepted as an ICML 2026 Spotlight paper.
- Core Optimization: Introduces program-aware routing at the scheduler level to address severe KV cache thrashing during agentic inference.
- Performance: Delivers 1.94x rollout throughput and 1.39x full-step throughput in matched synchronous GRPO runs. Achieves 2.5x higher single-node throughput and 10x lower P50 latency under high concurrency.
- Integration: The solution has been merged into the verl-recipe open-source project as a Dynamo recipe, enabling efficient multi-turn Agent RL training.
Related event: Together AI Unveils ThunderAgent, Doubling Agent Inference Throughput(7 posts)→
More from coding & agent
- LeanHEBO reimplements Huawei's algorithm 3x faster — hbouammar · 2026-08-25
- Agent runs autonomously for 24 days: System control beats pure model power — nodo48 · 2026-08-25
- Bananastand: CLI Tool to Check Real-time Value of RAM and Storage — dbreunig · 2026-08-25
- One Prompt Moves Your Coding Agent Session Across Claude, Codex, Pi and More — xhluca · 2026-08-25
- session-migrate: move coding agent sessions across Claude Code, Codex and more — xhluca · 2026-08-25
- Graph Engineering Guide: Building Road Maintenance AI Agents — MaryamMiradi · 2026-08-25