OpenVDN H3 live video generation promising on one GPU, but Ref2VA falls short in tests
eesahe · reddit · 2026-09-07
A user tested OpenVDN H3 — a live T2VA/I2VA/FL2VA video generation work that runs on a single GPU — finding the concept promising, though all official prompt examples are text2img and the work isn't trained for Ref2VA (reference-driven generation).
In the ComfyUI implementation (ComfyUI-VDN-H3), character model sheets were picked up well, but a specific expression reference (guruguru-me spiral eyes) couldn't be reproduced despite varying seeds, even though the base model can generate it. The author shared a detail comparison and hopes future versions train for Ref2VA.
More from Multimodal
- Why MiniMax H3 videos run long: the frames % 17 == 5 grid, explained — bennash · 2026-09-07
- Flux.2 Klein 9B identity drift: ReferenceLatent breaks down in action scenes — RecordEmbarrassed787 · 2026-09-07
- Structured Motion Graphics Prompt Template: Generate Pro Videos End-to-End — aziz4ai · 2026-09-07
- Beeble AI VFX test: ComfyUI still wins for material tracking — ssoissoi · 2026-09-07
- Dev builds a playground to test Gemini's music theory and tutoring skills — pitaru · 2026-09-07
- HeyGen open-sources hyperframes: write HTML, render video, built for AI agents — heygen-com · 2026-09-07