NVIDIA details HSTU recommender inference stack with up to 5.93x lower latency

PyTorch · x · 2026-10-09

Generative recommenders reframe personalization as sequence modeling over user behavior. NVIDIA's recsys-examples ships an end-to-end HSTU inference workflow with Dynamo-Triton: PyTorch AOTI compiles HSTU ranking models to native C++ artifacts, while FlexKV-backed KV caching stores reusable attention state to avoid recomputing long user histories. On an RTX PRO 6000 Blackwell GPU, an eight-layer HSTU model achieved up to 5.93x lower latency at batch size 8 with 100% KV-cache hit rate versus the same AOTI config without caching.

Original post →

More from Infra

Infra channel →