Audio-to-MIDI Model Released
petewoodbridge · x · 2026-07-10
MireloAI and Kyutai Labs have released an Audio-to-MIDI model.
It can take a complete recording as input, identify the instruments, and output track-separated MIDI results, including vocals, drums, bass, keyboards, etc. Unlike many existing solutions, this model does not require pre-separated stems and can process full mixed tracks directly.
Additionally, it detects chords, key, and tempo, providing producers with more musical context. The author noted that further explanations regarding the model, problem definition, and implementation methods have been published in the article.
Related event: Kyutai Releases MuScriptor for Multi-Instrument Transcription(4 posts)→
More from Multimodal
- Getting Started with AI Video: Solving Consistency and Censorship — cynicalnewenglander · 2026-07-22
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22