EvolvingAvatar Uses Test-Time Training to Make 3D Talking Heads Adapt as Conversations Unfold
HFUT-AI · hf · 2026-09-29
Interactive 3D head generation needs coordinated speaking/listening motion that evolves with the conversation, yet existing generators keep parameters fixed. EvolvingAvatar is a causal generator applying test-time training to adapt to user face video and dyadic audio during interaction.
- Dyadic context prediction provides a self-supervised signal without test-time motion labels
- Persistent fast weights accumulate in-conversation updates; transient jaw adaptation follows current audiovisual context, with predicted speech activity gating persistent adaptation
- Introduces InterHead-Bench, a 455.95-hour benchmark from single/dual-view conversation videos
- On the hardest OOD split, mismatch with recorded user-avatar expression statistics drops up to 11.1% as conversations unfold
More from Multimodal
- Midjourney v8.2 prompt: silver gelatin darkroom-style morning black-and-white portrait — tisch_eins · 2026-09-29
- One prompt, no cuts: Kling 4.0 video realism is getting hard to distinguish from real footage — FinanceYF5 · 2026-09-29
- One prompt, no cuts: Kling 4.0 video realism is getting hard to distinguish from real footage — FinanceYF5 · 2026-09-29
- A cyberpunk city built entirely in code: thousands of towers, live-synthesized soundtrack, zero audio files — techartist_ · 2026-09-29
- Atlases Are Already Inside: Recovering Population Templates by Making Diffusion Models Collapse — kwangmoo_yi · 2026-09-29
- SciGen-Verifier brings explainable, reasoning-driven verification to scientific image generation — Jiali Chen · 2026-09-29