STEMMA paper adds multi-audio song-to-stem reasoning to large audio-language models
keunwoochoi · x · 2026-10-10
A new arXiv paper introduces STEMMA, a multi-audio music QA framework for large audio-language models (LALMs).
- Built around production provenance: whether excerpts come from the same track/section, and which stems belong to which mixtures — relations existing LALMs and music QA datasets ignore.
- Uses a relation-first construction: define a target relation, then sample positive examples and hard negatives from the catalog; labels come directly from catalog provenance, not LLM-generated metadata.
- Ships STEMMA-Bench for evaluation and a track-disjoint training set, STEMMA-Instruct.
- Fine-tuning two LALMs on STEMMA-Instruct improves multi-audio reasoning, with the largest gains on catalog-determined structural relations, while preserving single-audio understanding.
More from Multimodal
- Vivix W1 real-time interactive video costs $0.003/sec, about 1/77th the price of Seedance clips — Scobleizer · 2026-10-10
- A 20-second wildfire video prompt built on one continuous ascending camera shot — umesh_ai · 2026-10-10
- Uncensored Qwen-Image Turbo GGUF hits Hugging Face trending — AtomicChat · 2026-10-10
- Do Anime Image Models Still Need Tag Soup in 2026? One Test Says Barely — Remarkable-Aspect879 · 2026-10-10
- SAM 3.1 + DINOv3 tracks surgical tools and tells apart near-identical twins on one A100 — MaziyarPanahi · 2026-10-10
- AI video editing is actually good now with your own footage — audrow · 2026-10-10