Marigold V2: single-step DiT depth estimation trainable on one 32GB consumer GPU
_akhaliq · x · 2026-09-10
Marigold V2 (to appear at SIGGRAPH Asia 2026) upgrades monocular depth estimation to a diffusion transformer: single-step inference, sharp edges, no OOM at 2K. Website, code, weights, and demo are all public.
- Built for small labs: starts from a big pretrained DiT (Qwen-Image-Edit-2509), 4-bit quantized with a rank-128 QLoRA, fine-tuned on a single 32GB consumer GPU — sensible depth in hours, done in days; no 80GB cards needed.
- Unconventional choices: V1 pushed depth maps through an image VAE; V2 aligns DiT features to ground-truth depth encoded via DINOv3 (iREPA-depth).
- Data insight: even the cleanest synthetic data (Hypersim) has broken depth at grass-blade scale, motivating SinkLoss, a loss tolerant of noisy ground truth without pixel-to-pixel correspondence.
- A 90-second paper walkthrough video is available.
Related event: Marigold V2: A Single-GPU DiT Depth Estimator Hits SIGGRAPH Asia 2026(11 posts)→
More from Research
- Steerable Visual Representations Presented as ICML Long Oral — y_m_asano · 2026-09-11
- OpenCVL: a satellite-to-photo registration dataset at ECCV 2026 — ducha_aiki · 2026-09-11
- Diverse VPR work submitted to ECCV 2026 — ducha_aiki · 2026-09-11
- EASE: evidence-anchored spatial attention lifts multimodal RLVR by up to 3.1 points, EMNLP 2026 — jiqizhixin · 2026-09-11
- P=NP Explained: Why Class Schedules and Circuit Routing Are the Real Hard Problems — thesaraharminta · 2026-09-11
- Hypothesis: ASI Has a Mathematical Incentive to Preserve Human Diversity — No_Cause_2731 · 2026-09-11