Fine-tune Qwen-Image-Edit into a depth estimator on one 32GB GPU with 4-bit QLoRA

AntonObukhov1 · x · 2026-09-09

Marigold V2 author Anton Obukhov shared the low-cost recipe:

Sensible depth after a few hours, done in days — no 80 GB card, no 8-GPU node. Inference is single-step and won't OOM at 2K resolution. Paper, code, weights, and demo links are all provided.

A textbook workflow for repurposing large pretrained models into specialized perception models on consumer hardware.

Related event: Marigold V2 Released: Single-GPU Fine-tuned DiT for Depth Estimation(5 posts)→

Original post →

More from Multimodal

Multimodal channel →