MiniMax-H3 Hybrid Model Trends on HF: Supports Text-to-Video & Audio-Video

smhfacct · hf · 2026-08-23

The model smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models is trending on Hugging Face. It is a hybrid model based on the Diffusion Transformer architecture, supporting text-to-video, image-to-video, and audio-video generation pipelines.

Original post →

More from Multimodal

Multimodal channel →