Hierarchical Self-Supervised World Models Enhance Music Co-Creation Agents
drscotthawley · x · 2026-08-06
Scott H. Hawley published the paper Helping Music Co-Creation Agents 'Listen' Well, introducing a Swin-Transformer masked-autoencoder (representation autoencoder). Trained on piano-roll images of MIDI, the model relies exclusively on internal-consistency objectives (such as LeJEPA attraction and cross-level masked-embedding prediction) with no reconstruction loss.
Experiments show that the learned six-level hierarchy effectively encodes musical structures:
- Chords & Keys: With a light auxiliary chord loss, chord root linear decoding accuracy jumps from chance to 90% on held-out songs.
- Phrase Boundaries: Boundary info increases monotonically toward coarser levels, peaking at 3× the pixel baseline at the top level.
- Equivariance: Near-perfect equivariance to pitch/time shifts exactly at the levels where geometric objectives act.
The author also released the STORMBIRD probe suite to evaluate nine encoder variants. Code and weights are expected to open-source around NeurIPS time.
Related event: MIDI-RAE-JEPA Enhances Music Agent Comprehension(3 posts)→
More from Multimodal
- Formas AI showcases 'Stone Edge' AI video generation demo — carlosbannon · 2026-08-06
- ComfyUI Showcases AI Post-Production Workflow with Custom LoRAs and Motion Nodes — petewoodbridge · 2026-08-06
- MiniMax H3 Music Video Production Tutorial and Workflow Released — petewoodbridge · 2026-08-06
- Hugging Face Showcases Cadena: Decompiling 3D Meshes into Editable CAD — huggingface · 2026-08-06
- Creator Uses Minimax Video Model to Generate Full Music Video — mementomori2344323 · 2026-08-06
- MiniMax H3 Tops Three Video Generation Categories, Beating ByteDance and Google — petewoodbridge · 2026-08-06