MLPerf Adds DLRMv4: HSTU Sequence Modeling Meets 560GB Embeddings for Production Recommenders
TheKanter · x · 2026-10-02
MLCommons has released DLRMv4, the next-generation recommendation training benchmark, developed by a task force from AMD, Meta, NVIDIA, and ByteDance.
- It replaces the old feature-interaction stack with HSTU, a transducer that reads a user's interaction history directly as a sequence, keeping the training benchmark aligned with the existing HSTU-based inference benchmark.
- It ships with a new dataset, Yambda-5B: a real multi-behavior music listening collection with 4.79B real interactions and a production-class embedding footprint of 560 GB.
- Context: hyperscale recommenders have shifted from aggregated dense features to per-event token sequences (driven by LLM-era sequence modeling), which keep improving with compute and parameters while older designs flatten out — the training suite previously had no workload covering this fast-growing category.
More from Infra
- Micron earnings show HBM still dominates AI memory, bit growth solid through 2028 — AccBalanced · 2026-10-02
- OpenRouter launches Security Center after finding 1,000+ dormant API keys across 85 employees — AccBalanced · 2026-10-02
- Morgan Stanley: Meta won't buy new chips to scale its Muse agent — AccBalanced · 2026-10-02
- Cerebras insiders dump stock as shares plunge $18 in a day, no word on lost GPT-6.1 deal — firstadopter · 2026-10-02
- Data Center Bottleneck Isn't GPUs: Transformer Lead Times Hit 115 Weeks — AccBalanced · 2026-10-02
- Cerebras-Linked Team Launches Detailed Series Explaining Disaggregated Inference — AccBalanced · 2026-10-02