PyTorchCon to Feature vLLM Sessions on Serving Stack
PyTorch · x · 2026-08-31
PyTorchCon NA 2026 will feature sessions on vLLM covering the serving stack, including attention and KV cache management, disaggregated serving, and hardware portability. Speakers from Red Hat, Amazon, IBM, NVIDIA, Google, Meta, Huawei, Mistral AI, and others will discuss expert parallelism and deployment across TPU, Trainium, Arm, and IBM Spyre.
More from Infra
- Model pricing cannot be inferred from total parameter count — xeophon · 2026-08-31
- Access to Nvidia GB300 is now open — francoisfleuret · 2026-08-31
- Beware of Neoclouds' security flaws: Container escapes and more — StasBekman · 2026-08-31
- LFS pause breaks AI training; configurable progress bar coming soon — janusch_patas · 2026-08-31
- Optimal Settings for Llama.cpp + Qwen 3.8: n-max 4 Fastest — GodComplecs · 2026-08-31
- AI designs chip from spec to hardware in 2 weeks — rohanpaul_ai · 2026-08-30