vLLM on K8s: GPU Faults Can Render a Node Unusable — Inference Is Stateful
tianyin_xu · x · 2026-10-09
Developer MadhavJivrajani posted a PSA: when running vLLM in production on Kubernetes, certain GPU faults can make a node effectively unusable — inference is a very stateful workload, even without the KV cache. The fix landed in vLLM main as PR #57635.
The PR fixes an autotune cache deadlock: flashinferautotune persists only rank 0's cache and broadcasts it on next engine start, but FlashInfer keys persisted MoE entries by tprank/eprank, so only rank 0's lookups hit. A cache hit skips the synchronized tuning pass's reduce, leaving ranks 1..N-1 blocked in the reduce while rank 0 spins at the trailing barrier until the 30-minute timeout. Triggers include a container restart with the cache dir on a pod-lifetime volume, or any second engine start seeing the file on persistent storage.
More from coding & agent
- Veteran dev: AI-coded it, I never read the code, but it's not vibe coding — judgment still matters — mjuric · 2026-10-09
- shadcn: the most important coding skill in the AI era is reading, not summarizing — shadcn · 2026-10-09
- Philosopher OKF: open-source LLM skill turns any topic into a structured study page — holyshitthatsucks · 2026-10-09
- Codex throttled to 5 tok/s as dev argues local model deployment is the only fix — lxfater · 2026-10-09
- Harrison Chase: trajectory labeling is several questions, not one pass/fail — Jev lands in LangSmith evals — hwchase17 · 2026-10-09
- Dev recreates seven Adobe apps in Rust with Claude, reigniting software copyright debate — technollama · 2026-10-09