Fine-tune Qwen-Image-Edit into a depth estimator on one 32GB GPU with 4-bit QLoRA
AntonObukhov1 · x · 2026-09-09
Marigold V2 author Anton Obukhov shared the low-cost recipe:
- Start from a big pretrained DiT (Qwen-Image-Edit-2509)
- Quantize to 4-bit
- Add a rank-128 QLoRA
- Fine-tune on a single 32 GB consumer GPU
Sensible depth after a few hours, done in days — no 80 GB card, no 8-GPU node. Inference is single-step and won't OOM at 2K resolution. Paper, code, weights, and demo links are all provided.
A textbook workflow for repurposing large pretrained models into specialized perception models on consumer hardware.
Related event: Marigold V2 Released: Single-GPU Fine-tuned DiT for Depth Estimation(5 posts)→
More from Multimodal
- Seedance 2.5 lands in CapCut desktop with 30-second generations and bulk editing — future_coded · 2026-09-09
- Combining two Seedance 2.5 prompt techniques for continuous shots and multi-cut sequences — techhalla · 2026-09-09
- 100 rule-based verifiers and 300 tasks: team builds a working RL recipe for video models — DanielKhashabi · 2026-09-09
- H3 MAX as a rendering engine? Demo promised with the right prompting — gorkem · 2026-09-09
- Gradium launches Voice Design: prompt-to-voice generation, free in API and Studio — ThePeterMick · 2026-09-09
- Making a self-correcting walking robot in Spline with GPT 6 Astra — dunkhippo33 · 2026-09-09