Paper proposes Ontology-based Target Sound Extraction for hierarchical audio separation
carlosheroliv · x · 2026-09-02
- Limitation: Current Target Sound Extraction (TSE) systems rely on fixed class representations, struggling with the hierarchical nature of environmental sounds.
- Proposed Task: Introduces "Ontology-based TSE," enabling a single model to extract sounds queried at any level of a sound ontology (e.g., fine-grained "cat" vs. high-level "animal").
- Method: Uses a learnable class embedding table over all nodes of an AudioSet-derived ontology, regularized with a Cophenetic Correlation Coefficient (CPCC) loss to align embedding distances with ontology tree distances.
- Result: Experiments demonstrate the benefits of incorporating ontology structure into TSE training.
More from Research
- AITHYRA opens fully funded PhD call for AI × biomedical research in Vienna — LucaAmb · 2026-09-02
- Yoav Goldberg: Don't write clickbait paper titles that strip key information — LChoshen · 2026-09-02
- Harness Arena: Open-Source Blind Benchmark Pits Claude Code, Codex and Other Agent Harnesses Head-to-Head — Due_Armadillo_8744 · 2026-09-02
- 'Posttraining is translation' partially retracted: verifiers can pull mass beyond the base model — Liu_eroteme · 2026-09-02
- H3 Acceleration Arena Trends on Hugging Face — multimodalart · 2026-09-02
- Self-hosted LLM with SLERP-merged GRPO experts outperforms larger baseline, serving half of production traffic — t-tech · 2026-09-02