Serve models from KitOps ModelKit on HAMi: a registry-native path to SGLang inference

HowDevelop · x · 2026-09-27

A demo of an end-to-end model serving pipeline: package a model as a versioned KitOps ModelKit (an OCI artifact), pull it into a Pod via a KitOps initContainer, schedule controlled GPU shares with HAMi, and serve it with SGLang's OpenAI-compatible API (optional vLLM co-location). The flow is registry-native — models are versioned on an OCI registry (Jozu Hub in the examples) — giving one clear path from model artifact to working inference on Kubernetes with NVIDIA GPUs.

Related event: KitOps ModelKit and HAMi Enable End-to-End GPU Inference Pipeline(2 posts)→

Original post →

More from coding & agent

coding & agent channel →