JD Open-Sources JoyAI-EchoWM, a World Model with Coherent Audio-Video Navigation

jiqizhixin · x · 2026-09-19

JD Explore Academy unveiled and open-sourced JoyAI-EchoWM at JDD 2026, an interactive audio-visual world generation model.

Current world models produce stunning video but break when users actually step inside: change your route or camera angle and footsteps, ambient sound, and voices stop staying coherent with the visuals. First-person and third-person controls are fundamentally different, and any gap shatters immersion.

EchoWM unifies control across first-person (camera moves with the observer) and third-person (camera follows subjects with motion coordinated to composition) perspectives, and generates synchronized audio—footsteps, environment sounds, speech—that remains coherent as the user navigates. Built on JoyAI-Echo's native audio-video generation, it creates a truly enterable, controllable, explorable world.

Original post →

More from Multimodal

Multimodal channel →