ThunderAgent by Together AI: 2x Faster Agentic Inference (ICML 2026)
togethercompute · x · 2026-07-30
Together AI introduced ThunderAgent, a high-throughput system for large-scale agentic inference, accepted as a Spotlight paper at ICML 2026.
The system introduces a novel program-level scheduler abstraction that effectively eliminates KV cache thrashing during multi-turn tool calls.
Key Performance Results:
- More than 2x single-node throughput, with roughly 10x lower P50 latency at high concurrency.
- Delivers a 2.4x speedup on an 8-node cluster, demonstrating near-linear scaling from 16 to 64 GPUs.
Engineering Compatibility:
- Designed as a drop-in solution: requires only a single programid field to plug into existing engine configs (such as KV offloading and speculative decoding).
- Format-agnostic by design, supports OpenAI chat completions today, and has already been adopted by SkyRL and NVIDIA Dynamo.
Related event: Together AI Launches ThunderAgent for 2x Faster Agent Inference(5 posts)→
More from Infra
- Buildcleaner reclaims 443GB of disk space by cleaning build artifacts, free and open-source MIT — jasonkneen · 2026-07-30
- LLM Inference Costs Drop Below $3 with B200s, Yet API Prices Stay High — AccBalanced · 2026-07-30
- Yann LeCun and Others Discuss: LLMs are the New Compilers, Performance is a Function of Compute — yisongyue · 2026-07-30
- Goldman Sachs Predicts 70x Jump in Monthly AI Token Processing by 2030 — Beth_Kindig · 2026-07-30
- Cold Start Benchmark: 244 GiB Model Loads in 154 Seconds — QuixiAI · 2026-07-30
- Unified FP8 in Training and Rollout Speeds Up RL by 16% — joecole · 2026-07-30