Helping Music Co-Creation Agents 'Listen': Hierarchical Self-Supervised World Models
Scott H. Hawley · hf · 2026-08-07
This paper presents a hierarchical self-supervised "world model" for symbolic music, designed to help collaborative music agents understand and generate music effectively. It uses a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives, requiring no labels or music-theory vocabulary.
The model spontaneously learns temporal and phrase structure, though chord recognition requires minimal supervision. Paired with a conditional flow-matching model, the system generates music suggestions in 2.8s on CPU (0.6s on Apple MPS) and supports graphical prompting via masked inpainting, offering a robust AI core for human-centric music creation.
More from Multimodal
- Underwater ink bloom portrait prompt: suspended figures in flowing pigment — aziz4ai · 2026-08-26
- Testing Qwen3.8 Vision: SVG Reconstruction & Anti-Benchmaxxing — bonobomaster · 2026-08-26
- LTX-2.5 Multishot Lip Sync on 8GB VRAM — big-boss_97 · 2026-08-26
- SREF Parameter Creates Anxious Animation Style with Muted Colors — tisch_eins · 2026-08-26
- CyberAgent Releases VAE Speech Align for Unsupervised Phoneme Alignment — kastnerkyle · 2026-08-26
- Qwen Create generates cinematic steampunk detective montage with Wan 3.0 — socialwithaayan · 2026-08-26