Training Image Editing Models Without Human Labels via Video Deltas
haremlifegame · reddit · 2026-08-08
A community discussion reveals a highly inspiring training method that can be used to fine-tune MiniMax H3 or build video model derivatives.
The core idea is to generate massive training data without human instructions:
- Take a video dataset and use a vision-language model (like Qwen) to analyze the video, specifically describing what changed between the first and last frames.
- Train the model to take the first frame and the 'description of changes' as input to generate the last frame.
- By inverting the order of operations, the model learns to execute instructions on 'what ought to change,' effectively becoming a powerful image editor.
This technique is an excellent approach for developing lighting LoRAs, image editing models, or fine-tuning video models.
More from Multimodal
- KIRI Open-Sources 3DGS Tool: Turning Static Scans into Interactive Animations — jonstephens85 · 2026-08-08
- fal Launches Agent: A Creative Assistant Integrating Image, Video, and 3D Models — gorkem · 2026-08-08
- Minimax H3 Turbo LoRAs Face Off: A 10-Scene Comparison — JoNike · 2026-08-08
- Ostris Teases Upcoming LoRA Weights for MiniMax H3 Model — krigeta1 · 2026-08-08
- Seedance 2.5 video-to-video test: 720p results show promise — mrjonfinger · 2026-08-08
- Looking for Workflows to Upscale MiniMax H3 Videos Using LTX 2.3 — Existing_Earth9000 · 2026-08-08