A²-Edit Enables High-Quality Image Editing with Coarse Masks via Expert Routing
机器之心 · wechat · 2026-07-31
Shanghai Jiao Tong University and Shanghai Institute of AI for Science introduced A²-Edit, a framework designed to overcome the difficulties of cross-category unified modeling and the reliance on precise masks in reference-guided image editing.
Core Designs:
- Hybrid Transformer Expert Routing: Extends conditional routing to jointly model Attention and FFN, dynamically selecting expert branches for different object categories (e.g., clothing, portraits, architecture).
- Mask Annealing Training Strategy (MATS): Progressively relaxes mask precision across training stages, shifting the model from relying on strict boundaries to contextual semantic reasoning. Ultimately, simple scribbles or bounding boxes suffice for editing.
- UniEdit-500K Dataset: A custom dataset containing 500k image pairs across 8 major categories and 209 sub-categories to break data homogenization.
Performance: Achieves top results in quantitative metrics and a 24-participant user study. The model shows almost no performance degradation when switching from precise to coarse masks. Limitations include high VRAM requirements (42GB) and potential ambiguity with oversized masks and no text prompts.
More from Multimodal
- Fish Audio Raises $52M Seed, Launches S2.1 Pro Voice Model — alexcovo_eth · 2026-07-31
- Daily Midjourney SREF: Graphic Novel Interiors with Hand-Drawn Shadows — tisch_eins · 2026-07-31
- Evolution of AI Video: Seedance 2.5 Coming to Higgsfield — socialwithaayan · 2026-07-31
- AI Visual Art: Blooming 'Blossoms of Flames' with striking VFX — TheChuckTone · 2026-07-31
- Seedance 2.5 Demo: High-Quality Short Film 'GOLIATH' — mementomori2344323 · 2026-07-31
- Testing AI Video Models on Complex Human Movement — HeyAmit_ · 2026-07-31