Light-MER: Sub-1B Parameter Model Outperforms 8B Teacher in Multimodal Emotion Recognition

新智元 · wechat · 2026-08-01

While multimodal LLMs excel at emotion recognition, their massive parameter counts (often 7B+) make them impractical for edge devices like robots. Researchers from the University of Glasgow, Shandong University, and CAS introduced Light-MER (accepted to ACMMM 2026), compressing an 8B teacher model into an 855M student model.

The methodology relies on two core components:

Experiments show that Light-MER achieves an average score of 74.61 across nine datasets, outperforming the 8B teacher model (73.93). Parameters and FLOPs are reduced by 11x, and peak VRAM drops from 20GB to 2.5GB, enabling high-quality, sub-second emotion recognition on consumer and edge hardware.

Original post →

More from Models

Models channel →