A production runbook for in-house LLM inference on Kubernetes
GD-Champ · reddit · 2026-07-29
Running in-house LLM inference on Kubernetes: a production runbook
A Reddit user shares a practical runbook for operating in-house LLM inference on Kubernetes, written while building the infra for their organization.
The post is framed as a production-oriented guide rather than a toy demo: it focuses on the engineering details needed to keep model serving workable in real deployments, and invites feedback from others who have built similar stacks.
More from Infra
- YouTube-style semantic IDs tackle recommender memory walls with dual-purpose tokens — _reachsumit · 2026-07-29
- Meta’s Memory Layer pushes Instagram Reels item coverage to 100% and freshness to 20 seconds — _reachsumit · 2026-07-29
- VaLiDRec uses variable-length LLM-aligned IDs and runs 87.49× faster than LC-Rec — _reachsumit · 2026-07-29
- India’s AI inference market will reward companies that co-optimize models and hardware — santoshpanda · 2026-07-29
- Hermes Agent Desktop impresses users with parallel tools and remote local-model setup — Teknium · 2026-07-29
- Chinese threat actor pivots infrastructure and leaks 775 API-key IDs from an AI reseller — cyb3rops · 2026-07-29