A modality-agnostic recommender learns one tokenizer from image and text data

_reachsumit · x · 2026-08-04

A modality-agnostic generative recommender learns one tokenizer for images and text

This paper proposes a modality-agnostic generative recommendation method that can learn from paired, image-only, and text-only item data.

Original post →

More from Multimodal

Multimodal channel →