Muka Robotics' LJM Model Ranks 2nd Globally in Embodied World Model Arena

机器之心 · wechat · 2026-07-30

Startup Muka Robotics, founded less than four months ago, proposed the Latent Joint-conditional Model (LJM). Trained with only 32 GPUs, it ranked 2nd globally on the authoritative WorldArena benchmark and achieved SOTA on four visual metrics of the LIBERO benchmark.

LJM uses a "dual-brain" architecture separating physical interaction reasoning from future video rendering. A reasoning expert predicts specific physical changes in the latent space, while a world modeling expert renders the video. This solves the flaw of traditional video models prioritizing visuals over physical interaction. Experiments prove that strong video priors alone are insufficient for robotic world models.

Original post →

More from Embodied

Embodied channel →