StemFX frames music mixing as token prediction on 105K songs

affige_yang · x · 2026-08-04

What the paper does

The authors describe StemFX, a music-mixing framework that treats mixing style as an autoregressive token prediction problem. It predicts variable-length FX chains for each source-separated stem using a Transformer decoder.

Key technical points

Results

On mixing style retrieval, StemFX beats all baselines across all tested chain lengths. On paired mixing style transfer, it achieves the best spectral fidelity and overall performance among compared methods.

Related event: StemFX Models Music Mixing Style as Token Prediction(2 posts)→

Original post →

More from Multimodal

Multimodal channel →