MIDI-RAE-JEPA: Hierarchical Representation Learning for Symbolic Music
drscotthawley · x · 2026-08-06
The author shared a new preprint, MIDI-RAE-JEPA, submitted to the NeurIPS Creative AI Track. The research presents a hierarchical representation learning approach for symbolic music encoded as piano roll images.
The method combines pitch- and time-shift equivariance objectives with LeJEPA and a Swin Transformer V2 encoder, trained entirely on self-supervised objectives. Experiments show that a decoder trained on frozen encoder embeddings achieves a reconstruction F1 of 0.995. Additionally, a flow matching generative model conditioned on these embeddings produces music that closely matches the pitch register and rhythmic density of the conditioning excerpt.
Related event: MIDI-RAE-JEPA Enhances Music Agent Comprehension(3 posts)→
More from Multimodal
- Testing Minimax H3 Image-to-Video with Default ComfyUI Workflow — JamesFilmsYT · 2026-08-06
- TwelveLabs Exec: Native Video Understanding is AI's Next Frontier — bigdata · 2026-08-06
- Porting Sana Cache Boosts MiniMax H3 Video Generation Speed by 1.39x — CeFurkan · 2026-08-06
- Hands-on with ByteDance's Seedance 2.5: Usable AI video in 1-2 tries, API coming soon — nikola_mr64990 · 2026-08-06
- MiniMax H3 Test: Accurately Generates Retro ASCII Art Cat Video — Perfect-Campaign9551 · 2026-08-06
- Creating an AI Music Video "Fuzzy Wuzzy" Using MiniMax — Peemore · 2026-08-06