Helping Music Co-Creation Agents 'Listen': Hierarchical Self-Supervised World Models

Scott H. Hawley · hf · 2026-08-07

This paper presents a hierarchical self-supervised "world model" for symbolic music, designed to help collaborative music agents understand and generate music effectively. It uses a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives, requiring no labels or music-theory vocabulary.

The model spontaneously learns temporal and phrase structure, though chord recognition requires minimal supervision. Paired with a conditional flow-matching model, the system generates music suggestions in 2.8s on CPU (0.6s on Apple MPS) and supports graphical prompting via masked inpainting, offering a robust AI core for human-centric music creation.

Related event: MIDI-RAE-JEPA: Hierarchical Self-Supervised World Model Helps Music Agents 'Listen'(5 posts)→

Original post →

More from Multimodal

Multimodal channel →