DIY Minimax H3 workflow adds multi-image + audio + video reference inputs
TheNeonGrid · reddit · 2026-09-24
The author built a custom workflow for Minimax H3 supporting multi-reference input: multiple images plus audio and/or video-audio references. Testing confirmed images + audio + video-audio combos work; video frames as reference combined with images didn't behave as hoped. The node dynamically creates more image/audio input sockets as needed; the workflow file is shared and better alternatives are welcome.
More from Multimodal
- Artificial Analysis launches TTS leaderboard with new Pronunciation Robustness Benchmark across 95 models — ArtificialAnlys · 2026-09-24
- Google Gemini 3.8 Flash TTS heard in examples: shorthand expansion and contextual pronunciation — ArtificialAnlys · 2026-09-24
- Sample audio: Gemini 3.8 Flash TTS contextually appropriate pronunciation demos — ArtificialAnlys · 2026-09-24
- Gemini 3.8 Flash TTS tops Artificial Analysis pronunciation benchmark at 89.5% — ArtificialAnlys · 2026-09-24
- Runway CEO declares video the final interface alongside real-time video UI demo — _AustinCalvert_ · 2026-09-24
- ACTx486 turns real podcast video into interactive synthetic conversation — karinanguyen · 2026-09-24