PyTorch AOTI Backend Speeds Up NVIDIA HSTU Inference by 1.14x–1.28x

PyTorch · x · 2026-09-04

PyTorch announces its AOTI backend delivers 1.14x–1.28x speedup over Python backend in NVIDIA's HSTU inference tests with Triton Inference Server. With KV cache, it achieves 2.20x–2.38x speedup in ideal all-GPU cache-hit scenarios. Results from NVIDIA's recsys-examples repo, showcasing best practices for generative recommenders on PyTorch.

Original post →

More from Infra

Infra channel →