vLLM on K8s: GPU Faults Can Render a Node Unusable — Inference Is Stateful

tianyin_xu · x · 2026-10-09

Developer MadhavJivrajani posted a PSA: when running vLLM in production on Kubernetes, certain GPU faults can make a node effectively unusable — inference is a very stateful workload, even without the KV cache. The fix landed in vLLM main as PR #57635.

The PR fixes an autotune cache deadlock: flashinferautotune persists only rank 0's cache and broadcasts it on next engine start, but FlashInfer keys persisted MoE entries by tprank/eprank, so only rank 0's lookups hit. A cache hit skips the synchronized tuning pass's reduce, leaving ranks 1..N-1 blocked in the reduce while rank 0 spins at the trailing barrier until the 30-minute timeout. Triggers include a container restart with the cache dir on a pod-lifetime volume, or any second engine start seeing the file on persistent storage.

Original post →

More from coding & agent

coding & agent channel →