Puppeteer: Diffusion Model Generates Object-Grounded, Posture-Aware Co-Speech Gestures

Pickford · hf · 2026-09-10

Puppeteer is a diffusion-based co-speech gesture generation model that uses causal latent tokens and object geometry to produce temporally coherent, physically grounded gestures that interact naturally with objects in the scene.

Original post →

More from Research

Research channel →