A modality-agnostic recommender learns one tokenizer from image and text data
_reachsumit · x · 2026-08-04
A modality-agnostic generative recommender learns one tokenizer for images and text
This paper proposes a modality-agnostic generative recommendation method that can learn from paired, image-only, and text-only item data.
- The model builds a shared semantic-ID tokenizer across different modalities.
- The goal is to unify recommendation learning even when training data is incomplete or unpaired.
- The result is a generative recommender that can exploit heterogeneous item sources without requiring every example to have both image and text.
More from Multimodal
- MiniMax H3 Open Weights Released with ComfyUI Integration — petewoodbridge · 2026-08-04
- MiniMax Exec Reviews Hailuo AI's Growth, Announces Push for Open Source — VoidAsuka · 2026-08-04
- Creating Poster Animations with Hailuo AI: Prompts and Results — LudovicCreator · 2026-08-04
- Qwen3.8-Max Tested: Generates Photorealistic Bugatti Engine in Three.js — cedric_chee · 2026-08-04
- RTX 3060 12GB Test: Generates 10s Portrait Video Locally in 17 Minutes — merica420_69 · 2026-08-04
- No More Messy Wires: Visual Debugger Tool for ComfyUI Released — niknah · 2026-08-04