OpenVDN H3 live video generation promising on one GPU, but Ref2VA falls short in tests

eesahe · reddit · 2026-09-07

A user tested OpenVDN H3 — a live T2VA/I2VA/FL2VA video generation work that runs on a single GPU — finding the concept promising, though all official prompt examples are text2img and the work isn't trained for Ref2VA (reference-driven generation).

In the ComfyUI implementation (ComfyUI-VDN-H3), character model sheets were picked up well, but a specific expression reference (guruguru-me spiral eyes) couldn't be reproduced despite varying seeds, even though the base model can generate it. The author shared a detail comparison and hopes future versions train for Ref2VA.

Original post →

More from Multimodal

Multimodal channel →