A production runbook for in-house LLM inference on Kubernetes

GD-Champ · reddit · 2026-07-29

Running in-house LLM inference on Kubernetes: a production runbook

A Reddit user shares a practical runbook for operating in-house LLM inference on Kubernetes, written while building the infra for their organization.

The post is framed as a production-oriented guide rather than a toy demo: it focuses on the engineering details needed to keep model serving workable in real deployments, and invites feedback from others who have built similar stacks.

Original post →

More from Infra

Infra channel →