Dual-H200 fine-tune of Marigold V2 with 16-frame temporal attention aims to fix video depth flicker
AntonObukhov1 · x · 2026-09-24
In a reply to Marigold author Anton Obukhov, breakdownart reveals they are fine-tuning Marigold V2 across two H200s, adding 16-frame temporal attention and training on Spring to reduce flicker in video depth estimation.
Early video comparisons look promising; next steps include SinkLoss and a native 1080p finishing phase. A concrete community route toward temporally consistent video depth.
More from Multimodal
- ECCV 2026 paper: x0-prediction fixes inefficient diffusion in reconstruction-tuned RAE latent spaces — serrjoa · 2026-09-24
- Open-source video cloning skill turns reference videos into editable, locally rendered projects — oran_ge · 2026-09-24
- AI video cloning can one-shot replicate a 3-minute viral promo, with editable details — oran_ge · 2026-09-24
- Where are the BFL 3 video and MiniMax H3 Max open weights? Community asks fal — krigeta1 · 2026-09-24
- RecCAR closes reciprocal cross-attention gap in joint video diffusion models — barilan · 2026-09-24
- ComfyUI Prompt Composer adds RefMod and LoRA support for subject-based prompt orchestration — Francky_B · 2026-09-24