Survey finds inconsistent data processing practices in offline recommender evaluation
_reachsumit · x · 2026-09-29
An arXiv survey (2609.31696) by Alberto Mancino et al. systematically characterizes data processing for offline recommender evaluation.
- Covers the full data-centric pipeline: dataset selection, interaction representation, data preparation, multimodal feature extraction, and train/validation/test splitting.
- Spans collaborative, sequential, session-based, graph-based, knowledge-aware, context-aware, multimodal, federated, cross-domain, contrastive-learning, and LLM-based recommendation.
- Key finding: pre-training data processing decisions shape comparability and reproducibility yet remain under-scrutinized, with highly inconsistent practices.
- Introduces a unified framework and taxonomy distinguishing data preparation from multimodal representation extraction.
More from Research
- Triangle Splatting SLAM: Imperial College's ECCV 2026 dense RGB-D SLAM with on-the-fly mesh extraction — rsasaki0109 · 2026-09-30
- Manifold opens early access: robotics eval platform runs thousands of GPU-parallel rollouts in 30 mins — paigeinsf · 2026-09-30
- 1,000 AI agents discover new CRISPR-like system in virus DNA within 24 hours — CurieuxExplorer · 2026-09-30
- Explaining just 5% of token positions retains nearly all audit success across 4.7M explanations — aisilab · 2026-09-30
- NTU's Persistence Forcing hits FID 1.63 on ImageNet 256 by heterogeneous refinement in pixel-space DiTs — NanyangTechnologicalUniversity · 2026-09-30
- IBM's Q&D trains proactive agents to ask better questions, beating a 15x larger model — ibm · 2026-09-30