Alibaba and CUHK’s RynnWorld-4D predicts a robot scene’s future from one image
jiqizhixin · x · 2026-07-26
Alibaba and CUHK’s RynnWorld-4D predicts a robot scene’s future in 4D from one image and a prompt
Alibaba DAMO Academy and CUHK researchers introduce RynnWorld-4D, a model that takes a single 3D image plus a text instruction and generates future frames, depth, and motion flow in one unified step.
- The system builds a richer physics-aware representation of how objects will move over time.
- On real-world bimanual manipulation tasks, it reportedly outperforms prior models, especially where spatial precision and temporal coordination matter.
- The post links to the project page, paper, code, and a Jiqizhixin report.
More from Embodied
- A OnePlus 12 can run Z Image Turbo and Flux.2 locally, but a 320×320 image still takes minutes — sgcego · 2026-07-26
- Orca opens a shared stack for dexterous hand research and teleoperation — tensorqt · 2026-07-26
- Jensen Huang says Optimus could help create a multi-trillion-dollar humanoid robot industry — XFreeze · 2026-07-26
- Weekly AI paper roundup covers world models, embodied AI, and self-improving agents — _akhaliq · 2026-07-26
- Vivix A1 accepts voice, text and images mid-interaction, demo says — Scobleizer · 2026-07-26
- Vivix A1 demo shows synchronized speech, gaze, facial motion and body movement — Scobleizer · 2026-07-26