MiniMax H3 Multimodal Generation: Mix Video, Image, Text, and Audio Inputs

piotrbinkowski · x · 2026-08-02

MiniMax announced an update for multimodal generation with its H3 model.

Users can now provide multiple types of references in a single generation:

By enabling H3 Multimodal Reference, the model can integrate all these inputs simultaneously to create content.

Original post →

More from Multimodal

Multimodal channel →