AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod
AWS ML Blog · rss · 2026-10-09
The problem
Multiple teams inside an enterprise—LLM training, CV inference, architecture research—need shared access to expensive GPU clusters. Without a well-designed multi-tenant setup, organizations get uncontrolled resource consumption, weak isolation, no cost attribution, and heavy admin overhead.
The solution
Amazon SageMaker HyperPod is a purpose-built service for large-scale gen AI compute clusters, orchestrated by EKS or Slurm, with automatic node health monitoring, fault recovery, and lifecycle management. The post presents a multi-tenant reference architecture on HyperPod with EKS:
- Authentication: AWS IAM Identity Center as the centralized layer, federating external IdPs like Microsoft Entra ID or Okta; supports SSO into SageMaker Studio and CLI workflows (aws sso login, then kubectl)
- Workspaces: per-team SageMaker AI domains with team-specific execution roles and a tailored Studio GUI
- Isolation: Kubernetes namespaces for workload isolation; EKS access entries map IAM roles to namespace-scoped RBAC, so both Studio and CLI requests are confined to the team's namespace
- Fairness: HyperPod Task Governance handles compute quotas and scheduling priorities; HyperPod Observability provides dashboards
- Storage: FSx for Lustre/OpenZFS with per-team shared dirs and per-user homes, plus per-team S3 buckets governed by IAM roles
Outcome
The architecture isolates teams end-to-end from auth through workload execution while sharing GPU infrastructure efficiently, and enables namespace-level cost allocation and chargeback.
More from Infra
- $2,800 rig of 8x Radeon Pro V620 (256GB VRAM) hits 3000+ t/s prefill via custom vLLM fork — _TheWolfOfWalmart_ · 2026-10-09
- Coatue lays out new forms of financing for AI compute — _AustinCalvert_ · 2026-10-09
- Benchmarking 4 open decision models on one RTX 4090: Laya fastest, Lev most accurate at 13x the latency — Fun-Meaning-6474 · 2026-10-09
- fal Engineer on Video Speed: H3 Max Renders 15 Seconds of Video in 5 — OdinLovis · 2026-10-09
- Microsoft and NVIDIA Going Hard on Local AI as a Defense Against Frontier Labs — MatthewBerman · 2026-10-09
- Eron v1.4: $2.99 native iOS client for Ollama with zero-buffer streaming and HomeKit tools — RA2B_DIN · 2026-10-09