Google adds video reference input to multimodal models for character consistency
shlomifruchter · x · 2026-08-28
A user praised Google's new video reference feature, noting its significant impact on maintaining character consistency. The feature allows dropping up to three seconds of reference video into multimodal inputs to map movement, visual context, and ensure consistency across scenes.
More from Multimodal
- Cartesia Releases Sonic-3.6 TTS, Toping AI Audio Benchmarks After Architecture Rebuild — HazyResearch · 2026-08-28
- Google Launches Gemini Omni 1.1 Flash with Scene Extension and 4K Upscaling — GoogleAI · 2026-08-28
- Alibaba's Wan3.0-Video Available on DeepInfra with 1080P Audio/Video Support — gharik · 2026-08-28
- Unbound Loop Update: Import 3D Meshes and 2D Images as Starting Points — rms80 · 2026-08-28
- Delivery App Design Built with Claude and Refined by Humans — Tegadesigns · 2026-08-28
- Cartesia Sonic-3.6 tops Voice Arena US English TTS leaderboard — hey_abusiddik · 2026-08-28