jxnl's robust music transcription pipeline: Astra + spectrogram error correction

jxnlco · x · 2026-09-24

jxnl shares his music transcription workflow for learning: use Astra to transcribe audio, then have the model review the spectrogram for error correction, and finally check the music theory, producing very robust transcriptions.

In an actual correction example, the model spotted clear issues: some wide vibrato became extra chromatic notes; around 3:22 the pitch tracker briefly jumped to a lower piano note while the clarinet continued above it. It corrected those, simplified isolated sixteenth rests, and noted the quiet ending remains ambiguous.

Original post →

More from Multimodal

Multimodal channel →