StemFX frames music mixing as token prediction on 105K songs
affige_yang · x · 2026-08-04
What the paper does
The authors describe StemFX, a music-mixing framework that treats mixing style as an autoregressive token prediction problem. It predicts variable-length FX chains for each source-separated stem using a Transformer decoder.
Key technical points
- A band-split multi-band CNN encoder with FiLM conditioning captures per-stem spectral structure.
- To scale training, the team extracted pseudo-stems from about 105K songs via source separation.
- They also built MultiAFx, a toolkit that unifies 85 audio effects from 7 Python libraries.
Results
On mixing style retrieval, StemFX beats all baselines across all tested chain lengths. On paired mixing style transfer, it achieves the best spectral fidelity and overall performance among compared methods.
Related event: StemFX Models Music Mixing Style as Token Prediction(2 posts)→
More from Multimodal
- Open-sourced SARAS, an AI video platform that turns topics into full videos — sai_teja_ · 2026-08-04
- Video-analysis tool v0.5.1 adds direct AI analysis with Gemini, Kimi, OpenAI and Claude — sujingshen · 2026-08-04
- MM H3 local test on a 3090 shows 500–900 second renders and high heat — TensorTinkererTom · 2026-08-04
- MM H3 local test on a 3090 shows 500–900 second renders and high heat — TensorTinkererTom · 2026-08-04
- ChatGPT prompt workflow produces a 10-second Blender orbit animation at 768×768 and 24 fps — goodside · 2026-08-04
- MiniMax-H3 text-to-video GGUF weights trend on Hugging Face — realrebelai · 2026-08-04