vivo's ART breaks the pseudo-target ceiling in makeup transfer with real-image supervision

机器之心 · wechat · 2026-09-05

An ECCV 2026 paper from vivo Blue Image Lab, HIT and Nanjing University tackles why makeup transfer fails on complex looks (sequins, face paint, stickers).

The bottleneck: paired data physically can't exist, so the field relies on pseudo-targets generated by large editing models — but student models inherit every flaw, capping quality at the "pseudo-target ceiling."

Method: two-stage training. Stage I trains on pseudo-targets plus an auxiliary makeup-removal model. Stage II keeps the transfer output as a differentiable "makeup carrier," composites it onto a makeup-free version of the reference, and uses reconstruction error against the real reference image as the supervision signal. Optimal controlled-noise strength: 0.6.

MF2K dataset: first 2K-resolution makeup dataset — 8,573 images at 2048×2048 across bare/light/heavy/artistic categories.

Results: ART ranks first on MSimG across all four benchmarks (9.22 vs 8.43 best baseline on the artistic subset) and wins all three axes in 864 user blind ratings. Direct verification shows outputs exceed the pseudo-targets they trained on, even when generated by weaker models like StableMakeup.

Takeaway: a transferable recipe for editing tasks lacking paired data — bootstrap with synthetic data, supervise with real images. Supports native 2048×2048 output; limitation: no explicit lighting/material modeling.

Original post →

More from Multimodal

Multimodal channel →