A modality-agnostic recommender learns one tokenizer from image and text data
_reachsumit · x · 2026-08-04
A modality-agnostic generative recommender learns one tokenizer for images and text
This paper proposes a modality-agnostic generative recommendation method that can learn from paired, image-only, and text-only item data.
- The model builds a shared semantic-ID tokenizer across different modalities.
- The goal is to unify recommendation learning even when training data is incomplete or unpaired.
- The result is a generative recommender that can exploit heterogeneous item sources without requiring every example to have both image and text.
More from Multimodal
- Runway Big Pitch Contest entry 'ALL THE WATER' imagines oceans vanishing underground — bennash · 2026-09-21
- MiniMax H3 Video: How to Use a Second Image to Replace a Body Part — Next-Place0 · 2026-09-21
- Paper: Physically Based Rendering in the Latent Space — ssh4net · 2026-09-21
- Adaptive Color Grading paper: KNN beats end-to-end models at tonescale prediction — ssh4net · 2026-09-21
- Blender as Director, Seedance as Renderer: A Cinematic AI Video Workflow — CurieuxExplorer · 2026-09-21
- MiniMax H3 Video Model Passes Fast-Cut Montage Test With Consistent Characters — Hailuo_AI · 2026-09-21