PyTorchCon NA puts vLLM center stage: KV cache, disaggregated serving, expert parallelism
PyTorch · x · 2026-09-30
vLLM features throughout PyTorchCon North America: Simon Mo keynotes on scaling open frontier inference infrastructure (core architecture, KV cache management, GPU kernels), with sessions on attention, KV cache transfer, disaggregated serving, elastic expert parallelism, cross-hardware scaling, custom accelerators, Trainium and RDMA KV, plus a dev meet-and-greet and posters.
More from Infra
- LLM-42 Paper at SOSP 2026 Brings Deterministic LLM Inference via Verified Speculation — tianyin_xu · 2026-09-30
- Indonesia's AI Debate: Owning a Data Center Doesn't Mean Owning the Decisions — AryHHAry · 2026-09-30
- Tesla secures $30B credit line to accelerate AI infrastructure, robotaxis and chip manufacturing — Polymarket · 2026-09-30
- Musk: Starship to launch at least 300GW of AI compute per year, targeting self-growing Moon and Mars cities — XFreeze · 2026-09-30
- Jensen Huang: These aren't data centers anymore, they're 'super intelligence factories' — TinfoilTricorn · 2026-09-30
- Musk: Memphis AI data center caused local 'over-employment', may have doubled tax budget — XFreeze · 2026-09-30