PyTorch AOTI Backend Speeds Up NVIDIA HSTU Inference by 1.14x–1.28x
PyTorch · x · 2026-09-04
PyTorch announces its AOTI backend delivers 1.14x–1.28x speedup over Python backend in NVIDIA's HSTU inference tests with Triton Inference Server. With KV cache, it achieves 2.20x–2.38x speedup in ideal all-GPU cache-hit scenarios. Results from NVIDIA's recsys-examples repo, showcasing best practices for generative recommenders on PyTorch.
More from Infra
- Analyst: no model trained without NVIDIA has ever beaten one trained on NVIDIA — BenBajarin · 2026-09-04
- A 2023 desktop PC now resells above its purchase price amid hardware inflation — Darpinian · 2026-09-04
- OpenAI's Jalapeño: RTL freeze to tapeout in 9 months, serving ChatGPT 10 weeks after first silicon — thehiphopswami · 2026-09-04
- Lenovo Unveils Dozens of AI Devices at IFA, RTX Spark Laptop Runs 120B-Param Models Locally — 智东西 · 2026-09-04
- Astra reportedly trained on 100k GPUs at Stargate Texas in first $1B training run — ai · 2026-09-04
- Spark-2.5-4B runs on $250 Jetson Orin Nano: 128K context, 2048-needle test 99.8% pass — Puzzleheaded_Base302 · 2026-09-04