HeyGen trains avatar behaviors as separate LoRAs, then composes them at inference

HeyGen · x · 2026-07-25

HeyGen says it can control multiple avatar behaviors without a jointly labeled dataset by training each behavior — expressiveness, gaze, and camera — as separate LoRAs on single-attribute data.

At inference time, the behaviors are composed together, and the train/inference gap is narrowed with a short fine-tune on real overlapping samples.

Original post →

More from Multimodal

Multimodal channel →