SheetSage2 uses synthetic supervision to lead music transcription on 12 of 15 benchmarks

m-a-p · hf · 2026-10-08

m-a-p presents SheetSage2, a unified music-to-lead-sheet transcription framework combining synthetic supervision (auto-annotated MIDI rendered to audio), task-specific structured decoding for musically coherent output, and autoregressive distillation that removes task-specific dynamic programming at inference. A single SheetSage2-AR model beats listed prior systems on 12 of 15 benchmark-metric pairs across 8 collections, substantially improving on SheetSage1. Model weights and inference code are publicly available.

Original post →

More from Multimodal

Multimodal channel →