CutBCE: TPU Kernel Eliminates OOM in Large-Vocabulary Recommendation Training, 91.9% Faster

_reachsumit · x · 2026-10-06

A new paper introduces CutBCE, an exact, hardware-accelerated binary cross-entropy (BCE) loss and gradient operator built in JAX/Pallas for industrial sequential recommender systems with massive item catalogs (10^5–10^7 items).

The problem: Full-vocabulary BCE training materializes a dense [B, N, V] logits tensor in HBM, incurring O(BNV) memory and fatal OOM errors. While chunked Softmax loss optimizations exist for LLMs, large-scale multi-label BCE remained unexplored.

Techniques:

Results: Eliminates OOM with up to 91.9% speedup on single-chip TPU v5e/v6e; on 8-chip TPU training of multi-label SASRec with 876k items (Yambda-50M), peak HBM drops 65.7%.

Original post →

More from Infra

Infra channel →