Tutorial: Run SGLang on Kubernetes with HAMi GPU Shares — Memory Quotas and Compute Throttling

HowDevelop · x · 2026-09-04

A hands-on lab (Lab 15) shows how to install the HAMi scheduler on an existing NVIDIA GPU Kubernetes cluster and run SGLang inference with GPU shares: software-enforced VRAM quotas (nvidia.com/gpumem) and compute throttling (nvidia.com/gpucores) per Pod, cloud-vendor-agnostic. The result is an OpenAI-compatible API verified via /v1/models and /v1/chat/completions, with quota enforcement confirmed inside the Pod. Needs a GPU node with >25,000 MiB free VRAM (tested on H100 80GB); 45 minutes to reproduce.

Original post →

More from Infra

Infra channel →