PyTorch Replaces CUDA with FBTriton for Embedding Kernels: 1.28x Faster Forward, 2x Backward
PyTorch · x · 2026-10-07
Meta's team migrated Table Batched Embedding (TBE) kernels in PyTorch recommender systems from legacy CUDA to Triton (FBTriton), achieving up to 1.28x faster forward passes and 2x faster backward passes.
Key technical points
- TBE fuses embedding lookup and pooling across many tables into a single GPU launch, cutting launch overhead and improving memory efficiency
- Two forward implementations: a generic gather path and a fast path using a small-table histogram
- Tuned bags-per-program and gather width; int32 indices/offsets when ranges fit under 2^31, halving index storage
A deep engineering writeup with measured gains and future optimization directions — highly relevant for teams serving large-scale embeddings across sharded GPUs.
More from Infra
- OpenSBI and Linux now boot on Maxion cores of ET-SOC1 cards, porting done agentically — glenbeer · 2026-10-07
- Samsung's Q3 profit reportedly set to soar 770% YoY as AI demand overwhelms memory supply — Polymarket · 2026-10-07
- openTPU: A One-Person Open-Source AI Accelerator Designed by AI Agents, Running 10 LLMs on an FPGA — ai · 2026-10-07
- Free Online Conference All Day AI Set for Oct 22 With Talks on SLMs and Local Inference — FikoFox · 2026-10-07
- Best $4,000 Local LLM Rig? Weighing R9700s, Strix Halo, and 6x Arc B60 — BinaryGrind · 2026-10-07
- Paris' Tour Montparnasse Plans 120,000 Nvidia B300s Delivering 540 EFLOPS at 304MW — IgorCarron · 2026-10-07