Tutorial: Run SGLang on Kubernetes with HAMi GPU Shares — Memory Quotas and Compute Throttling
HowDevelop · x · 2026-09-04
A hands-on lab (Lab 15) shows how to install the HAMi scheduler on an existing NVIDIA GPU Kubernetes cluster and run SGLang inference with GPU shares: software-enforced VRAM quotas (nvidia.com/gpumem) and compute throttling (nvidia.com/gpucores) per Pod, cloud-vendor-agnostic. The result is an OpenAI-compatible API verified via /v1/models and /v1/chat/completions, with quota enforcement confirmed inside the Pod. Needs a GPU node with >25,000 MiB free VRAM (tested on H100 80GB); 45 minutes to reproduce.
More from Infra
- 12B local models now rival GPT-4, and tiered local AI is coming, argues Wardell — draginol · 2026-09-04
- QTS's own site confirms zero-water cooling at its data centers — GlenBradley · 2026-09-04
- Virginia State Study: Most Data Centers Use No More Water Than a Large Office Building — GlenBradley · 2026-09-04
- Inference platform Lighter is 'absolutely exploding,' August update shows — econoar · 2026-09-04
- NVIDIA-backed Open Secure AI Alliance moves to the Linux Foundation — CackleRooster · 2026-09-04
- AI data centers could lower local electricity rates by spreading fixed grid costs, expert explains — alexvoica · 2026-09-04