Marigold V2 Hits New SOTA in Monocular Depth Estimation with Single-Step Diffusion Transformers
AntonObukhov1 · x · 2026-09-11
Marigold V2 (arXiv:2609.08084) repurposes diffusion transformers (DiT) into state-of-the-art monocular depth estimators, setting a new SOTA on Sintel surface normal estimation.
Key techniques:
- Single-step inference from pretrained multi-step flow-matching models, with optional quantization — preserving capacity while staying cheap to run
- Two fixes for naive-training artifacts: aligning internal representations with semantic features from ground truth, and a two-stage fine-tuning protocol built on a novel Sinkhorn-based loss
- Sharper, cleaner depth maps with strong out-of-distribution generalization; 16-26% AbsRel improvement over the previous best on KITTI and ETH3D
- Resolves fur, foliage, and hair-thin edges that eluded prior models
Authors are from EPFL, University of Bologna and others; evals are live on Papers with Code.
More from Multimodal
- Synthesia Adds Custom Avatars Built From a Text Prompt — synthesiaIO · 2026-09-11
- CapCut brings Seedance 2.5 to desktop with 30-second generations and GPT Image 2 workflow — anthara_ai · 2026-09-11
- A Dancer Balancing on a Ladder in Purple Light: A Striking AI Video Demo — misovalko · 2026-09-11
- Arabic-Supported AI Video Models Mapped, With a Lip-Sync Workflow Using Gemini TTS and Magnific — aziz4ai · 2026-09-11
- Recreating a Minimax song with YuE2 via a Hermes agent in pure CLI on a 16GB GPU — wzwowzw0002 · 2026-09-11
- Users complain Imagine 2.5 outputs look stiffer and blander than before — Valuable-Butterfly65 · 2026-09-11