Pushing H3 text-to-video with layered prose prompts that Qwen3VL actually understands

SIR_NVAX_A_LOT · reddit · 2026-09-06

A Redditor shares an H3 text-to-video experiment stress-testing how well Qwen3VL handles non-standard, prose-heavy prompts — with surprisingly good comprehension.

Key techniques: a layering approach composing the shot through long, cascading prose that defines the subject in detail (face material, luminous hair filaments, carved star-chart wardrobe, lighting); int8 quantization at 20 steps, 720p; quick test-runs at 2s/int8/8 steps. The prose style adds randomness, so composition and lighting sometimes took a third attempt. The final piece stitches four clips together, with the score also generated by H3 using a 32x32 canvas trick. A full example prompt (an extremely detailed layered description of an "artificial divinity" character) is included.

Original post →

More from Multimodal

Multimodal channel →