Marigold V2: diffusion transformer depth estimation trained on one GPU in days

AntonObukhov1 · x · 2026-09-09

Marigold V2 (to appear at SIGGRAPH Asia 2026) is out. V1 post-trained an image generator into a depth estimator on a single GPU — the accessible research game; V2 upgrades to a diffusion transformer, producing very sharp edges in depth maps and normals, and is quite versatile.

Collaborator Matteo Poggi notes the whole thing trains in a few days on one GPU, no cluster needed, keeping the low-barrier, reproducible spirit alive.

Related event: Marigold V2: Single-GPU DiT Depth Estimator Hits Five Benchmarks(8 posts)→

Original post →

More from Multimodal

Multimodal channel →