Pushing H3 text-to-video with layered prose prompts that Qwen3VL actually understands
SIR_NVAX_A_LOT · reddit · 2026-09-06
A Redditor shares an H3 text-to-video experiment stress-testing how well Qwen3VL handles non-standard, prose-heavy prompts — with surprisingly good comprehension.
Key techniques: a layering approach composing the shot through long, cascading prose that defines the subject in detail (face material, luminous hair filaments, carved star-chart wardrobe, lighting); int8 quantization at 20 steps, 720p; quick test-runs at 2s/int8/8 steps. The prose style adds randomness, so composition and lighting sometimes took a third attempt. The final piece stitches four clips together, with the score also generated by H3 using a 32x32 canvas trick. A full example prompt (an extremely detailed layered description of an "artificial divinity" character) is included.
More from Multimodal
- 12 best AI image generators in 2026 compared, with no universal winner — TheTuringPost · 2026-09-06
- Turn industrial CAD files into interactive sales demos with Astra's 3D capabilities — VibeMarketer_ · 2026-09-06
- Creator shares WIP short film trailer "ONE SMALL BITE" made with Seedance — Resident_Subject2213 · 2026-09-06
- GPT-6 Astra-generated 3D game demoed on a 5070 Ti, with Mac/Windows clients packaged for download — op7418 · 2026-09-06
- Testing Hailuo on a 15-second prehistoric action scene: cliff jump onto a flying pterosaur — azed_ai · 2026-09-06
- Inside an AI film shot list: a 27-second video broken into dozens of second-long cuts — techhalla · 2026-09-06