ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation
togethercompute · x · 2026-07-30
Together AI introduced ThunderAgent, a high-throughput system designed for large-scale agentic inference. By introducing a novel program abstraction for scheduling agentic LLM requests, it effectively eliminates KV cache thrashing.
Key Performance Results:
- Achieves >2× single-node throughput with roughly 10× lower P50 latency at high concurrency.
- Delivers a 2.4× speedup on an 8-node cluster, demonstrating near-linear throughput scaling from 16 to 64 GPUs.
Engineering Integration:
- Acts as a drop-in solution requiring only a single programid field, plugging into existing engine configs like KV offloading and speculative decoding.
- Format-agnostic by design, supports OpenAI chat completions today, and has already been adopted by SkyRL and NVIDIA Dynamo.
The research was accepted as a Spotlight paper at ICML 2026.
Related event: Together AI Unveils ThunderAgent for 2x Agent Inference Speedup(6 posts)→
More from Infra
- Musk Reveals xAI Infrastructure: Minihard and Macroharder Pack 220k GB300s — kevinnbass · 2026-07-30
- Down $600M in a Day: Inside Leopold's AI Infrastructure Investment Thesis — ivan_bezdomny · 2026-07-30
- YC Paper Club Dives into Multi-GPU Kernels, Inference Efficiency and Heterogeneous Hardware — Y Combinator · 2026-07-30
- SK Hynix Earnings Analysis: AI Memory Demand Strong, Market Overreacts to Oversupply — tengyanAI · 2026-07-30
- Single 8x 5090 Rig Hits 167k tokens/s Training Throughput, Beating DDP — jon_durbin · 2026-07-30
- Buildcleaner reclaims 443GB of disk space by cleaning build artifacts, free and open-source MIT — jasonkneen · 2026-07-30