ImIR replaces text prompts with image instructions for all-in-one restoration
Süleyman Aslan · hf · 2026-09-23
ImIR proposes replacing text prompts with "image instructions" for all-in-one image restoration.
- Prior recipes adapt a large pretrained image-editing model with a LoRA adapter plus a text prompt; ImIR instead derives a continuous instruction vector from the degraded image itself
- The image enters via two paths: structure through the VAE, and semantics through a lightweight token mapper that shifts the degraded image's vision-language embedding toward a clean image's embedding
- Because the instruction is continuous, scaling it yields a family of valid restorations for non-unique targets like low-light enhancement
- One adapter trained in 3 hours on a single GPU adapts a Qwen-Image-Edit model to six tasks
- Matched comparisons show image instructions beat text conditioning, and enable task-agnostic restoration without degradation labels
More from Multimodal
- PixVerse R2 hands-on: AI video that becomes a world you walk through with WASD — HeyAmit_ · 2026-09-23
- One-Person AI Music Video Took a Month: 20 Stills and 10 Takes Per Clip — kraussian · 2026-09-23
- MiniMax H3 Ref2VA Freezes at Model Initializing With an 8th Reference Image — itchplease · 2026-09-23
- H3 long-video degradation workaround: a second noise-injection refine stage in latent space — xyzdist · 2026-09-23
- Dev builds a character-swap LoRA dataset end-to-end with Codex and GPT Image 2.5 — ostrisai · 2026-09-23
- Tencent Hunyuan Image3.5 preview went live Sept 22, free for a limited time — HeyAmit_ · 2026-09-23