Reddit user benchmarks audio models for music-video prompts; Qwen3.5 Omni Plus wins
mwoody450 · reddit · 2026-09-13
A Reddit user tested audio-capable models for a music-video workflow: feed a 3:17 track (anonymized as song.mp3 to prevent name cheating), have models produce timestamped structural descriptions, then feed text to a stronger model for video generation. Results:
- Qwen3.5 Omni Plus: best overall — exact track length, nailed the 1:04 orchestral buildup and 1:17 chorus surge, solid thematic summary
- Gemini 3.8 Flash: richest output but timestamps ran a few seconds early
- Mimo 2.5 Thinking: good tempo-shift detection and correct 3:15 fade identification, missed the first buildup
- Gemini 3.1 Pro Preview: missed the 1:17 beat drop
- Muse Spark 1.2/1.3 and Inkling Thinking never actually received the audio and hallucinated entirely fictional songs
More from Multimodal
- Seedream team unveils VoT: visual thinking before pixel rendering for image generation — JingxiangSun42 · 2026-09-13
- Wan 2.2 Suddenly Outputs Wavy Abstract Video; Blackwell fp16 Instability Suspected — deviruchii · 2026-09-13
- Krea 2 Recognizes Most Source Characters Without LoRAs, User Finds — magik_koopa990 · 2026-09-13
- Tencent's Open-Source WeMM-Embedding Runs Video Retrieval on an RTX 3060 — huangyun_122 · 2026-09-13
- Artist Generates Retro-Futuristic Terraforming GIFs with Impasto Oil Painting Texture Using Codex and Astra — mhmazur · 2026-09-13
- MiniMax Generates a Short Video Draft on 8GB VRAM in 17 Minutes — Creative_aidumpster · 2026-09-13