Together AI teams with IBM and NVIDIA for dedicated B300 inference cluster on IBM Cloud
togethercompute · x · 2026-10-08
- Together AI is partnering with IBM and NVIDIA to scale enterprise inference, starting with a dedicated NVIDIA B300 GPU cluster on IBM Cloud with Spectrum-X Ethernet networking — the first dedicated large-scale inference cluster of its kind on IBM Cloud, with Together as the first customer.
- Division of labor: NVIDIA supplies silicon and networking, IBM provides the cloud, Together runs the inference layer.
- Together now serves hundreds of trillions of tokens per month to over a million developers, and demand keeps climbing.
- The pitch: enterprises pick open models for data sovereignty plus frontier-level performance at a fraction of closed-model cost. IBM Cloud GM Alan Peacock: spending on tokens is worth it if the business benefits.
Related event: Together AI Partners with IBM and NVIDIA on B300 Inference Clusters(2 posts)→
More from Infra
- Blockway Ships Agens Volundr 32B: Only 18 Layers Keep KV Cache for Long Context — lmoroney · 2026-10-08
- Yaroslav Bulatov Traces the Origins of the Memory Wall in New Research Writeup — yaroslavvb · 2026-10-08
- Two Years Building a Local Agent OS: Multi-Node Routing, Memory Arbitration and Failure Governance — Budget_One_8784 · 2026-10-08
- DARPA QBI finalists revealed: six quantum approaches, up to $300M each through 2029 — hardimanjames · 2026-10-08
- DCVC recalls backing Atom Computing from pitch deck to DARPA QBI finalist — hardimanjames · 2026-10-08
- Rust+Vulkan Inference Backend Goes Cross-Vendor: Pick Kernels by CPU ISA, Not GPU Vendor — PhysicsDisastrous462 · 2026-10-08