RINO: Unifies Vision Tasks as RGB Editing
Timing Yang · hf · 2026-07-15
This work introduces RINO, which unifies various visual information into RGB image representations and reformats vision tasks into an RGB-to-RGB image editing problem.
Core Concept
- Different visual signals like natural images, masks, and depth maps are all processed using a single RGB encoding/decoding architecture.
- Different tasks share the exact same set of parameters, eliminating the need for task-specific heads or models.
- This approach is analogous to how language models process text: covering diverse tasks through a unified interface.
Results
- Based on a general image editing backbone, with no task-specific fine-tuning.
- Demonstrates robust zero-shot performance in dense understanding tasks like segmentation and depth estimation.
- Also shows competitive performance in dense-conditioned generation tasks such as pose-to-image generation.
The authors hope this will inspire unified vision-language systems, allowing different visual tasks to be expressed and solved using a shared "visual language." The code is open-source.
More from Multimodal
- Skywork AI Video packages storyboarding, editing, generation and export into one workspace — _jaydeepkarale · 2026-07-21
- WiseMe turns voice replies into text, images, files, and demo videos from your own knowledge — JaynitMakwana · 2026-07-21
- Reddit shares an AI-generated mini movie called The Lunar Ship — Ermajean12 · 2026-07-21
- AI creator GossipGoblin is turning short-form clips into a feature film — Hackedv12 · 2026-07-21
- TimeLens2 claims SOTA on 7 video grounding benchmarks with 4B and 8B models — _akhaliq · 2026-07-21
- AI-made 4-minute horror short ‘THE NOT KNOW’ lands as a shareable demo — gen_ericai · 2026-07-21