H3 T2VA Test: Layered Prompts Outperform R2VA

SIR_NVAX_A_LOT · reddit · 2026-08-21

A user shared their experience with the video generation model H3. The author believes that T2VA (Text to Video with Audio) mode, especially with layered prompts, is actually stronger than R2VA. However, for maintaining character consistency using character sheets, FL2VA+R2VA remains the go-to choice, and its voice cloning capability is top-notch. The post showcases content generated using T2VA (bf16/50 steps).

Original post →

More from Multimodal

Multimodal channel →