NoSpoon music-video automation still breaks because models cannot really hear music
Kyrannio · x · 2026-07-23
The author is testing an automated workflow for creating and editing music videos with NoSpoon, but says it still fails on a basic requirement: the models cannot really “hear” music.
They note that beat mapping only goes so far, and sung lyrics can throw the system off badly. In their view, the workflow will remain unreliable until models understand the nuances of listening to music much better.
More from Multimodal
- Reddit user shares an img2vid workflow built on an RTX 4070 Super — Professional_Wash169 · 2026-07-23
- TwelveLabs shows a video memory layer that stores moments, entities and themes — AI Engineer · 2026-07-23
- Early Krea 2 training results show the model already producing visuals — darlens13 · 2026-07-23
- ChatGPT Image 2.0 turns out a moody, classical-style painting — DeryaTR_ · 2026-07-23
- ChatGPT Work turns hundreds of Slack photos into a montage during a commute — gabrielchua · 2026-07-23
- Team built a demo video in two days with Remotion, Claude, and OpenAI vision — bosmeny · 2026-07-23