AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod

AWS ML Blog · rss · 2026-10-09

The problem

Multiple teams inside an enterprise—LLM training, CV inference, architecture research—need shared access to expensive GPU clusters. Without a well-designed multi-tenant setup, organizations get uncontrolled resource consumption, weak isolation, no cost attribution, and heavy admin overhead.

The solution

Amazon SageMaker HyperPod is a purpose-built service for large-scale gen AI compute clusters, orchestrated by EKS or Slurm, with automatic node health monitoring, fault recovery, and lifecycle management. The post presents a multi-tenant reference architecture on HyperPod with EKS:

Outcome

The architecture isolates teams end-to-end from auth through workload execution while sharing GPU infrastructure efficiently, and enables namespace-level cost allocation and chargeback.

Original post →

More from Infra

Infra channel →