MODUS turns a decoder-only model into a single any-to-any multimodal system

EPFL-VILAB · hf · 2026-07-29

MODUS proposes a decoder-only any-to-any framework that can predict any modality from any combination of others in a single model.

Original post →

More from Multimodal

Multimodal channel →