ByteDance's TM20K: Knowledge Distillation Framework for Long-Sequence Ad Recommendation
_reachsumit · x · 2026-08-10
ByteDance recently published new research on ultra-long behavior sequence modeling for e-commerce ad recommendation. Traditional methods face low training efficiency and throughput bottlenecks when scaling sequence length, often at the cost of fine-grained information.
The paper introduces TM20K, balancing effectiveness and efficiency via full transformer modeling combined with a two-stage knowledge distillation framework. Specifically, a heavily trained teacher model processes full sequence tokens, while well-motivated token merge approaches are designed for student models to significantly compress sequence length while maintaining performance. This approach successfully scales ad sequences from 5K to 20K.
More from Research
- GPT-6 Astra claims Terminal-Bench Science lead at 65.7%, 31.4 points clear of second place — DeryaTR_ · 2026-09-21
- Ant International's Code2Skill mines verifiable agent skills from code at scale — ant-intl · 2026-09-21
- RecreationWorld: five-platform CUA benchmark — GPT-6 Astra hits 58.1% but passes all tests on just 2.8% — Shuai Bai · 2026-09-21
- OmniVBench: a 12k-checklist benchmark and 340K-sample dataset for omni reference-to-video generation — Wenxue Li · 2026-09-21
- AI reviews training AI reviewers: study finds 'scientific-judgment collapse' and an open-source fix — Sy-Tuyen Ho · 2026-09-21
- Apple's MintAct unifies GUI agents across mobile, desktop, and web, hitting SOTA 48.9 on OSWorld-Verified — apple · 2026-09-21