Marigold V2 retools diffusion transformers for sharper monocular depth estimation
huawei-bayerlab · hf · 2026-09-09
Huawei Bayer Lab released Marigold V2, repurposing diffusion transformers (DiTs) for monocular depth estimation. Key ingredients:
- Single-step flow-matching inference
- Semantic alignment
- A Sinkhorn-based two-stage fine-tuning protocol
The result is noticeably sharper out-of-distribution depth maps plus strong results on related dense regression tasks.
Related event: Marigold V2: Single-GPU DiT Depth Estimator Accepted to SIGGRAPH Asia 2026(6 posts)→
More from Multimodal
- Saudi dev crafts ChatGPT prompt for low-credit motion graphics, doable from a phone — aziz4ai · 2026-09-09
- Full text-to-video prompt template for raw found-footage alien planet vlogs — techhalla · 2026-09-09
- VoiceStudio open-sources a local ElevenLabs-style stack: voice clone, dubbing, 646 languages — thisguyknowsai · 2026-09-09
- OpenAI image-2.5 hands-on: sub-$0.01, 23s generations clearly beat Muse — tobowers · 2026-09-09
- GPT image model generates a 10-second GIF, sparking video-model takeover talk — rounak · 2026-09-09
- TapNow accuses influencer of stripping watermark to pass off signed artist's SD 2.5 AI film as rival product demo — TomLikesRobots · 2026-09-09