Audio-to-MIDI Model Released

petewoodbridge · x · 2026-07-10

MireloAI and Kyutai Labs have released an Audio-to-MIDI model.

It can take a complete recording as input, identify the instruments, and output track-separated MIDI results, including vocals, drums, bass, keyboards, etc. Unlike many existing solutions, this model does not require pre-separated stems and can process full mixed tracks directly.

Additionally, it detects chords, key, and tempo, providing producers with more musical context. The author noted that further explanations regarding the model, problem definition, and implementation methods have been published in the article.

Related event: Kyutai Releases MuScriptor for Multi-Instrument Transcription(4 posts)→

Original post →

More from Multimodal

Multimodal channel →