REALM generates reactive listener facial motion, deployed on an Ameca humanoid robot

MacquarieUni · hf · 2026-09-29

REALM is a coarse-to-fine framework for audio-driven reactive listening in embodied conversational AI. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio via delay-centered attention and adaptive gating; a coarse decoder predicts base motion, refined by audio-conditioned stochastic residuals for blinks and brief expressions. It outperforms baselines on ViCo and L2L, and was deployed on an Ameca humanoid with a perceptual user study. Code and demo are public.

Original post →

More from Embodied

Embodied channel →