MIDI-RAE-JEPA: Hierarchical Representation Learning for Symbolic Music

drscotthawley · x · 2026-08-06

The author shared a new preprint, MIDI-RAE-JEPA, submitted to the NeurIPS Creative AI Track. The research presents a hierarchical representation learning approach for symbolic music encoded as piano roll images.

The method combines pitch- and time-shift equivariance objectives with LeJEPA and a Swin Transformer V2 encoder, trained entirely on self-supervised objectives. Experiments show that a decoder trained on frozen encoder embeddings achieves a reconstruction F1 of 0.995. Additionally, a flow matching generative model conditioned on these embeddings produces music that closely matches the pitch register and rhythmic density of the conditioning excerpt.

Related event: MIDI-RAE-JEPA Enhances Music Agent Comprehension(3 posts)→

Original post →

More from Multimodal

Multimodal channel →