PyTorch Replaces CUDA with FBTriton for Embedding Kernels: 1.28x Faster Forward, 2x Backward

PyTorch · x · 2026-10-07

Meta's team migrated Table Batched Embedding (TBE) kernels in PyTorch recommender systems from legacy CUDA to Triton (FBTriton), achieving up to 1.28x faster forward passes and 2x faster backward passes.

Key technical points

A deep engineering writeup with measured gains and future optimization directions — highly relevant for teams serving large-scale embeddings across sharded GPUs.

Original post →

More from Infra

Infra channel →