ByteDance's TM20K: Knowledge Distillation Framework for Long-Sequence Ad Recommendation

_reachsumit · x · 2026-08-10

ByteDance recently published new research on ultra-long behavior sequence modeling for e-commerce ad recommendation. Traditional methods face low training efficiency and throughput bottlenecks when scaling sequence length, often at the cost of fine-grained information.

The paper introduces TM20K, balancing effectiveness and efficiency via full transformer modeling combined with a two-stage knowledge distillation framework. Specifically, a heavily trained teacher model processes full sequence tokens, while well-motivated token merge approaches are designed for student models to significantly compress sequence length while maintaining performance. This approach successfully scales ad sequences from 5K to 20K.

Original post →

More from Research

Research channel →