MiniMax H3 experiment blends reference image and 10s clip into a seamless 12s video

Hailuo_AI · x · 2026-09-08

A Japanese creator tested MiniMax H3 by combining a reference image with a 10-second reference video to generate a coherent 12-second clip — an overhead view (reference image) transitioning to books seen through a window (reference video). The assets were made with other tools: the overhead image via GPT Image2, and the 3D reference video in Blender using GPT Astra.

Original post →

More from Multimodal

Multimodal channel →