SheetSage2 uses synthetic supervision to lead music transcription on 12 of 15 benchmarks
m-a-p · hf · 2026-10-08
m-a-p presents SheetSage2, a unified music-to-lead-sheet transcription framework combining synthetic supervision (auto-annotated MIDI rendered to audio), task-specific structured decoding for musically coherent output, and autoregressive distillation that removes task-specific dynamic programming at inference. A single SheetSage2-AR model beats listed prior systems on 12 of 15 benchmark-metric pairs across 8 collections, substantially improving on SheetSage1. Model weights and inference code are publicly available.
More from Multimodal
- Reddit user shoots a cinematic short of H.G. Wells' The Cone with MiniMax H3 — nikhilprasanth · 2026-10-08
- Ideogram ships 4.5 image model with new app UI and video integration — bennash · 2026-10-08
- Dreamina launches Dreamina Originals, first content label for AI films and series — LudovicCreator · 2026-10-08
- Reverse-Engineer Any Video into a Seedance Prompt with Claude or Gemini — TawohAwa · 2026-10-08
- MiniMax demos code-driven video: text to JS animation and explainer videos — Hailuo_AI · 2026-10-08
- Turkish TTS trained from scratch with the Drifting method on a single RTX 5090 — kadir_nar · 2026-10-08