Helping Music Agents 'Listen' Well: Hierarchical Self-Supervised World Models

drscotthawley · x · 2026-08-06

The author shared their latest preprint, Helping Music Co-Creation Agents 'Listen' Well, which presents a hierarchical self-supervised "world model" for symbolic music.

The core of the model is a 2.55M-parameter Swin V2 encoder trained with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), requiring no labels or music-theory vocabulary. Probing experiments reveal that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, while note density and harmonic detail off the finest.

Furthermore, by introducing a small chord-supervision head, joint chord recovery increased from 0.18 to 0.54, and key detection (which was never supervised) jumped from 0.16 to 0.70. Combined with conditional flow-matching, the model achieves a target window reconstruction F1 of 0.996 in pixel space.

Related event: MIDI-RAE-JEPA Enhances Music Agent Comprehension(3 posts)→

Original post →

More from Multimodal

Multimodal channel →