TriLayer: Explicit Video Layer Modeling for Realistic Object Insertion

postech-cglab · hf · 2026-07-30

Most current video editing systems lack explicit layered video representations, limiting their ability to perform realistic compositing, object reuse, and consistent manipulation. Existing methods rely on implicit inference or per-scene optimization.

To address this limitation, researchers introduced TriLayer, a large-scale triplet video dataset containing aligned composite, background, and foreground videos. The foreground layers include both object appearance and associated visual effects.

Building on this dataset, the authors proposed DBL-Diffusion, a dual-branch diffusion framework that jointly models RGB composites and RGBA foreground layers through shared denoising and cross-branch interaction. The framework is instantiated in two tasks:

Experiments demonstrate that explicit layer modeling substantially improves both insertion fidelity and decomposition quality.

Original post →

More from Multimodal

Multimodal channel →