Semantic subword tokenization improves generative recommenders by reducing intra-item attention overload
_reachsumit · x · 2026-08-25
The paper proposes Semantic Subword Tokenization (SST), which compresses user history into reusable subword tokens instead of fixed codes, reducing intra-item attention overload and freeing capacity to model item relationships. Experiments on three public datasets and three generative recommender backbones show improvements over fixed-length and variable-length SID baselines. Accepted to CIKM 2026.
More from Research
- AAAI 2027 review debate: Should empirical papers without code be auto-rejected? — SimpleObvious4048 · 2026-08-25
- Study asks: Are LLM agents time-aware and budget-conscious? — maksym_andr · 2026-08-25
- SA-RSQ: Sparse Representation Framework for Multi-modal Recommender Systems — _reachsumit · 2026-08-25
- Revisiting N2DCG: Empirical Reformulation for Carousel Recommendation — _reachsumit · 2026-08-25
- RAG collapse: LLM answers converge when retrieving self-authored content, 79.6% simulations collapse — _reachsumit · 2026-08-25
- Spotify study: better reasoning traces can hurt recommender effectiveness — _reachsumit · 2026-08-25