Light-MER: Sub-1B Parameter Model Outperforms 8B Teacher in Multimodal Emotion Recognition
新智元 · wechat · 2026-08-01
While multimodal LLMs excel at emotion recognition, their massive parameter counts (often 7B+) make them impractical for edge devices like robots. Researchers from the University of Glasgow, Shandong University, and CAS introduced Light-MER (accepted to ACMMM 2026), compressing an 8B teacher model into an 855M student model.
The methodology relies on two core components:
- SWD-H (Sliced Wasserstein Distance Hidden-state Distillation): Aligns the hidden state distributions between teacher and student, transferring the structural representation of multimodal emotional cues rather than just mimicking outputs.
- M-GRPO (Multi-Reward GRPO): Optimizes generation quality by rewarding emotion accuracy, appropriate length, and information density, effectively reducing verbose outputs and inference latency.
Experiments show that Light-MER achieves an average score of 74.61 across nine datasets, outperforming the 8B teacher model (73.93). Parameters and FLOPs are reduced by 11x, and peak VRAM drops from 20GB to 2.5GB, enabling high-quality, sub-second emotion recognition on consumer and edge hardware.
More from Models
- DeepSeek's Ultimate Philosophy: Maximizing Intelligence Throughput Per GPU-Second — teortaxesTex · 2026-08-01
- AI Market Irony: Just Lower Prices to Achieve the 'Pareto Frontier' — andersonbcdefg · 2026-08-01
- Anthropic Accused of Shifting Stance on Models' Reluctance to Be Deprecated — repligate · 2026-08-01
- Teknium Tests DeepSeek V4 Flash: Full Agent Task Costs Just $0.07 — Teknium · 2026-08-01
- Claude 4 Fails Long-Context Retrieval, Suspected KV Compression Artifacts — teortaxesTex · 2026-08-01
- OpenAI Offers Free GPT-5.6 to 100K Researchers as Harvard Physicist Cites 100x Speedup — 新智元 · 2026-08-01