vLLM and NVIDIA Co-host Meetup on Scaling LLM Inference Efficiency

vllm_project · x · 2026-08-12

The vLLM project is co-hosting a meetup with the NVIDIA Dynamo team in San Francisco on August 24.

The event will feature tech talks focused on serving LLMs efficiently at scale. Discussions will cover the latest work across vLLM and NVIDIA Dynamo, including inference optimization, distributed serving, and practical challenges of running these systems in production.

Original post →

More from Companies & People

Companies & People channel →