Marigold V2 released: DiT-based depth estimator trained on a single consumer GPU
AntonObukhov1 · x · 2026-09-09
Marigold V2 (to appear at SIGGRAPH Asia 2026) is out. V1 post-trained an image generator into a depth estimator on a single GPU; V2 upgrades to a diffusion transformer, producing very sharp edges and covering multiple modalities: depth, surface normals, and albedo. Paper, code, weights, and demo are all publicly available.
Related event: Marigold V2 Released: Single-GPU Fine-tuned DiT for Depth Estimation(5 posts)→
More from Multimodal
- Combining two Seedance 2.5 prompt techniques for continuous shots and multi-cut sequences — techhalla · 2026-09-09
- 100 rule-based verifiers and 300 tasks: team builds a working RL recipe for video models — DanielKhashabi · 2026-09-09
- H3 MAX as a rendering engine? Demo promised with the right prompting — gorkem · 2026-09-09
- Gradium launches Voice Design: prompt-to-voice generation, free in API and Studio — ThePeterMick · 2026-09-09
- Making a self-correcting walking robot in Spline with GPT 6 Astra — dunkhippo33 · 2026-09-09
- Comfy Clarifies MiniMax H3 Licensing: Free Community Tier, Enterprise-Priced Commercial — crystal_alpine · 2026-09-09