Hierarchical Self-Supervised World Models Enhance Music Co-Creation Agents

drscotthawley · x · 2026-08-06

Scott H. Hawley published the paper Helping Music Co-Creation Agents 'Listen' Well, introducing a Swin-Transformer masked-autoencoder (representation autoencoder). Trained on piano-roll images of MIDI, the model relies exclusively on internal-consistency objectives (such as LeJEPA attraction and cross-level masked-embedding prediction) with no reconstruction loss.

Experiments show that the learned six-level hierarchy effectively encodes musical structures:

The author also released the STORMBIRD probe suite to evaluate nine encoder variants. Code and weights are expected to open-source around NeurIPS time.

Related event: MIDI-RAE-JEPA Enhances Music Agent Comprehension(3 posts)→

Original post →

More from Multimodal

Multimodal channel →