Viggle-Animate open-weights: video character replacement from one repainted frame, no pose or mask pipeline
cocktailpeanut · x · 2026-09-05
- Viggle released its first open-weight model, Viggle-Animate, a 33.1B full finetune of MiniMax-H3's ref2va transformer, jointly distilled with DMD down to three forward passes — 5s of video in 26s on one GPU.
- Usage is minimal: repaint one frame of the driving video with a new character in any image editor; the model propagates the edit across the shot, leaving motion, camera and timing untouched.
- The key innovation is dropping the traditional pipeline entirely: no pose skeletons, segmentation masks, face tracking, identity encoders, or text prompts. Since the reference is a frame of the clip itself, pose, camera, framing and lighting already align with the footage.
- It's strongest where replacement is hardest: fast motion like whipping heads and full kicks, tracked frame by frame rather than smeared.
- Commenters highlight this as a workflow "unbundling" trend: offload pose detection and segmentation to an external image editor for a single frame, making the video model itself fast — a workflow innovation implemented as a finetune.
More from Multimodal
- Reddit user compares two MiniMax H3 Turbo LoRAs for video generation — No-Bee-231 · 2026-09-05
- MiniMax H3 workflow: REF2VA visuals + FL2VA audio + LightX2V speed, 198s on RTX 5090 — gabxav · 2026-09-05
- AI short film "Ishiro - The Last Blade" shared on Reddit — Nervous-Barnacle-888 · 2026-09-05
- Dev calls for a MiniMax H3 finetune that fills speech gaps with human filler words — cocktailpeanut · 2026-09-05
- Nvidia DLSS 5 frame interpolation discussed in Stable Diffusion community — KonoTheSavage1 · 2026-09-05
- Perceptron's Multilook API Prefills Video Context Once, Cuts Input Cost to 32% at 16 Prompts — AkshatS07 · 2026-09-05