Video Generation Test: ref2vid Loses Character Facial Details vs img2vid

Shadow-Games-1909 · reddit · 2026-08-10

A creator shared hands-on experience with a video generation model (likely H3). When generating long continuous scenes, the ref2vid mode handled character positioning better and reduced background mismatches compared to img2vid using first and last frames.

However, the author noted a significant quality drop in ref2vid, particularly the severe loss of character facial details. Even with a 4K upscaler, the visual quality falls far short of the original static image. The creator questions whether reference-based generation inevitably sacrifices this much image quality.

Related event: MiniMax H3 Tested: Ref2Vid vs I2V Quality Differences(2 posts)→

Original post →

More from Multimodal

Multimodal channel →