NoSpoon music-video automation still breaks because models cannot really hear music

Kyrannio · x · 2026-07-23

The author is testing an automated workflow for creating and editing music videos with NoSpoon, but says it still fails on a basic requirement: the models cannot really “hear” music.

They note that beat mapping only goes so far, and sung lyrics can throw the system off badly. In their view, the workflow will remain unreliable until models understand the nuances of listening to music much better.

Original post →

More from Multimodal

Multimodal channel →