PyTorchCon China: Viettel details a Kubernetes-based Token-as-a-Service inference platform

PyTorch · x · 2026-09-09

At PyTorchCon China 2026, Viettel's Trong Vinh Nguyen showed how PyTorch inference workloads run on Kubernetes: CNCF tooling for GPU pooling and multi-tenant slicing, PyTorch Foundation projects for serving optimization, OIF at the bare-metal layer, plus batch scheduling and GPU allocation across a mixed GPU fleet — yielding a Token-as-a-Service platform.

Related event: Viettel Unifies GPU Clusters with Open-Source Stack for Token-as-a-Service(2 posts)→

Original post →

More from Infra

Infra channel →