Reddit user benchmarks audio models for music-video prompts; Qwen3.5 Omni Plus wins

mwoody450 · reddit · 2026-09-13

A Reddit user tested audio-capable models for a music-video workflow: feed a 3:17 track (anonymized as song.mp3 to prevent name cheating), have models produce timestamped structural descriptions, then feed text to a stronger model for video generation. Results:

Original post →

More from Multimodal

Multimodal channel →