NVIDIA details HSTU recommender inference stack with up to 5.93x lower latency
PyTorch · x · 2026-10-09
Generative recommenders reframe personalization as sequence modeling over user behavior. NVIDIA's recsys-examples ships an end-to-end HSTU inference workflow with Dynamo-Triton: PyTorch AOTI compiles HSTU ranking models to native C++ artifacts, while FlexKV-backed KV caching stores reusable attention state to avoid recomputing long user histories. On an RTX PRO 6000 Blackwell GPU, an eight-layer HSTU model achieved up to 5.93x lower latency at batch size 8 with 100% KV-cache hit rate versus the same AOTI config without caching.
More from Infra
- Chollet: AI capex is growing super-exponentially while progress is only sub-linear — fchollet · 2026-10-09
- FineWeb author: annotating pretraining data with a 27B model is wild but pays off at deployment — antoine_chaffin · 2026-10-09
- Patching MLX to stage quantized weights to FP8 yields +40% prefill on M6 — Brilliant-Hall1387 · 2026-10-09
- a16z: Agents burn 5x the tokens of humans, up 14x in six months as AWS rewires for machine users — a16z · 2026-10-09
- Microsoft's Surface RTX dev box ditches ConnectX-7 NIC, limiting multi-box clustering — Scobleizer · 2026-10-09
- PyTorch now runs natively on AWS Trainium, with live demo at PyTorch Conference NA — PyTorch · 2026-10-09