Marigold V2 repurposes diffusion Transformers for dense prediction, trainable on a single GPU in days
RexDouglass · x · 2026-09-14
- Marigold V2 adapts diffusion Transformers for dense prediction by finetuning Qwen-Image-Edit-2509 with 4-bit quantization and Rank-128 QLoRA.
- Estimates depth, normals, albedo and more in a single step.
- Training runs on a single GPU within days; video applications (with STeR) hinted at.
More from Multimodal
- AI-Generated Short Film 'The Ninth Bell': A Heist Thriller in a Locking Opera House — Icy-Grand-4355 · 2026-09-14
- AI can now generate unlimited camera angles from a single video, including unseen views — TheMoonMidas · 2026-09-14
- Redditor gives ChatGPT total creative freedom to pack maximum detail into one image — NVDA808 · 2026-09-14
- Rigged Mesh Can Still Fail in Motion: Lessons from Three Chibi 3D Character Tests — softmarshmallow · 2026-09-14
- Comic-book heroes recreated with 'GPT-6 Astra', fal camera control, Three.js — OdinLovis · 2026-09-14
- Turning video games into 150,000-word audiobooks with Claude, GPT and Kokoro — HighSeasArchivist · 2026-09-14