STEER: steerable 3D head avatar motion prior for two-person conversation, code released

rsasaki0109 · x · 2026-09-17

STEER (SIGGRAPH Asia 2026) released inference code. It is a causal flow-matching motion prior generating a target person's 3D facial motion (61-D FLAME state: expression, jaw, neck, eye rotation, eyelids) in dyadic conversation, conditioned on the partner's motion/audio and the target's own audio, with explicit semantic control over gaze, head rhythm, and emotion. The repo includes the prior + autoencoder pipeline and FLAME visualization, outputting side-by-side comparison videos.

Original post →

More from Multimodal

Multimodal channel →