KAIST's AnyTalk: video diffusion models generate 3D speech animation for arbitrary characters

KAIST · hf · 2026-08-18

KAIST presents AnyTalk, which generates 3D speech animations for arbitrary characters without any animation data.

The approach: adapt video diffusion models via character-specific fine-tuning to synthesize talking-head videos, then optimize blendshape parameters from the synthesized footage to obtain drivable 3D speech animation. A distilled real-time variant is also provided.

Original post →

More from Multimodal

Multimodal channel →