PyTorchCon to Feature vLLM Sessions on Serving Stack

PyTorch · x · 2026-08-31

PyTorchCon NA 2026 will feature sessions on vLLM covering the serving stack, including attention and KV cache management, disaggregated serving, and hardware portability. Speakers from Red Hat, Amazon, IBM, NVIDIA, Google, Meta, Huawei, Mistral AI, and others will discuss expert parallelism and deployment across TPU, Trainium, Arm, and IBM Spyre.

Original post →

More from Infra

Infra channel →