Serve models from KitOps ModelKit on HAMi: a registry-native path to SGLang inference
HowDevelop · x · 2026-09-27
A demo of an end-to-end model serving pipeline: package a model as a versioned KitOps ModelKit (an OCI artifact), pull it into a Pod via a KitOps initContainer, schedule controlled GPU shares with HAMi, and serve it with SGLang's OpenAI-compatible API (optional vLLM co-location). The flow is registry-native — models are versioned on an OCI registry (Jozu Hub in the examples) — giving one clear path from model artifact to working inference on Kubernetes with NVIDIA GPUs.
Related event: KitOps ModelKit and HAMi Enable End-to-End GPU Inference Pipeline(2 posts)→
More from coding & agent
- PrimeScientist: MCTS-based policy teaches research agents where to spend their experiments — rohanpaul_ai · 2026-09-27
- PRIMESCIENTIST teaches research agents to allocate experiments, +10.3% reward with 50.6% fewer attempts — rohanpaul_ai · 2026-09-27
- Reviewer: Claude Opus 5.5's real win is taste, not benchmarks — usage far lower than expected — Hesamation · 2026-09-27
- dhh ports Omarchy screensaver from Rust to assembly with Opus, hits 450x speedup — AccBalanced · 2026-09-27
- A git worktree + meta-skill pattern for managing agent skills across multiple projects — JnBrymn · 2026-09-27
- FFS MCP Server lets you manage RingCentral feature flags via natural language — modelcontextprotocol · 2026-09-27