AMRD: Adaptive Multi-Teacher Distillation for Lightweight Speech Emotion Recognition

Yuqi Li · hf · 2026-07-31

While large self-supervised models excel at Speech Emotion Recognition (SER), their high compute cost makes them unsuitable for edge devices. This paper proposes Adaptive Multi-teacher Relational Distillation (AMRD) to compress these models into lightweight students.

AMRD tackles two key challenges: a one-class SVM dynamically evaluates teacher reliability per batch to assign weights, and a relational distillation loss aligns the inter-sample similarity matrices between teacher and student, preserving structural information missed by standard logit matching. On the IEMOCAP and CREMA-D datasets, AMRD outperforms single-teacher baselines in most settings, with ablations confirming the complementary gains of both components.

Original post →

More from Research

Research channel →