Together Compute unveils multi-node inference engine with near-linear scaling, adopted by SkyRL and NVIDIA Dynamo
togethercompute · x · 2026-07-30
Together Compute announced its inference engine achieves near-linear multi-node scaling: from 671 to 2,248 steps/min when scaling from 16 to 64 GPUs, with lead widening from 1.79x at 2 nodes to 2.39x at 8. The engine is drop-in compatible via a single programid field, supports KV offloading and speculative decoding, is format-agnostic with OpenAI chat completions support, and is already adopted by SkyRL and NVIDIA Dynamo.
Related event: Together AI Launches ThunderAgent for 2x Faster Agent Inference(5 posts)→
More from Infra
- Buildcleaner reclaims 443GB of disk space by cleaning build artifacts, free and open-source MIT — jasonkneen · 2026-07-30
- LLM Inference Costs Drop Below $3 with B200s, Yet API Prices Stay High — AccBalanced · 2026-07-30
- Yann LeCun and Others Discuss: LLMs are the New Compilers, Performance is a Function of Compute — yisongyue · 2026-07-30
- Goldman Sachs Predicts 70x Jump in Monthly AI Token Processing by 2030 — Beth_Kindig · 2026-07-30
- Cold Start Benchmark: 244 GiB Model Loads in 154 Seconds — QuixiAI · 2026-07-30
- Unified FP8 in Training and Rollout Speeds Up RL by 16% — joecole · 2026-07-30