EchoWM: Omnimodal World Model with 6-DoF Navigation Support

Songchun Zhang · hf · 2026-08-25

EchoWM is an open omnimodal world model capable of generating synchronized high-resolution video, sound, music, and speech following continuous 6-DoF navigation trajectories. It supports both first- and third-person views, enhancing multimodal generation for embodied AI.

Original post →

More from Multimodal

Multimodal channel →