REALM generates reactive listener facial motion, deployed on an Ameca humanoid robot
MacquarieUni · hf · 2026-09-29
REALM is a coarse-to-fine framework for audio-driven reactive listening in embodied conversational AI. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio via delay-centered attention and adaptive gating; a coarse decoder predicts base motion, refined by audio-conditioned stochastic residuals for blinks and brief expressions. It outperforms baselines on ViCo and L2L, and was deployed on an Ameca humanoid with a perceptual user study. Code and demo are public.
More from Embodied
- Leaked OpenAI 'dot' details show raising phone to ear triggers ChatGPT Voice — koltregaskes · 2026-09-29
- Uniformation S10 Robotic Resin Printer Handles Printing, Washing and Drying Automatically — philfung · 2026-09-29
- Korea's Diden Spider: a magnetic-legged welding robot that climbs shipyard walls — MickeySteamboat · 2026-09-29
- "Terminator" robot demoed at REK arena as man-vs-machine extreme sport takes shape — cixliv · 2026-09-29
- AIgirl demo shows a robot making dinner — miamiamorph · 2026-09-29
- HexaAnything represents robot tasks as code, beating VLA baselines on RoboCasa365 — Hongcheng Gao · 2026-09-29