Google adds video reference input to multimodal models for character consistency

shlomifruchter · x · 2026-08-28

A user praised Google's new video reference feature, noting its significant impact on maintaining character consistency. The feature allows dropping up to three seconds of reference video into multimodal inputs to map movement, visual context, and ensure consistency across scenes.

Original post →

More from Multimodal

Multimodal channel →