MODUS turns a decoder-only model into a single any-to-any multimodal system
EPFL-VILAB · hf · 2026-07-29
MODUS proposes a decoder-only any-to-any framework that can predict any modality from any combination of others in a single model.
- Unlike many existing any-to-any systems, it does not start from scratch with encoder-decoder or diffusion architectures.
- Instead, it reuses strong pre-trained decoder-only models as a prior and treats all modalities symmetrically.
- The model supports arbitrary modalities as both inputs and outputs without modality-specific heads, losses, or task pipelines.
- The paper highlights use cases such as chained generation through intermediate modalities and cross-modal self-verification by scoring one generated modality with another.
- MODUS reports strong out-of-the-box performance and competitive results against specialist and multitask baselines across benchmarks.
- All materials are open-sourced at the project site.
More from Multimodal
- Creator says an album was made end to end with AI using Suno — Kyrannio · 2026-07-29
- A layout guide for image generation that helps models stop placing elements randomly — sujingshen · 2026-07-29
- ComfyUI tutorial shows KREA 2 Identity Edit v1.2 for low-VRAM face editing — cgpixel23 · 2026-07-29
- GPT Image 2 storyboard packs a pilot’s tea break and a mecha battle into one 3x3 grid — Practical_Low29 · 2026-07-29
- Prompt tweak turns a generic scene into a Saudi heritage image with traditional details — aziz4ai · 2026-07-29
- Twelve Labs unveils a video intelligence stack built around search, memory and agentic workflows — qdrant_engine · 2026-07-29