ZimaBlue: Evolving Generalizable World Action Models via Video Pre-training
JoyFutureAcademy · hf · 2026-09-02
ZimaBlue learns generalizable world action models from large-scale egocentric video via a three-stage curriculum and a slow-fast architecture. It substantially improves zero-shot robotic manipulation by leveraging video data to understand physical world interactions.
More from Embodied
- Expert: 1 million humanoids in US jobs within a decade, maybe — binarybits · 2026-09-02
- Digimon AI Project: Reward Hacking and Progress in PPO Training — redfoxkiller · 2026-09-02
- Qwen-Drive-1.0: A Vision-Language Foundation Model for Autonomous Driving — Qwen · 2026-09-02
- Analyst report massively overestimates robot data generation — zephyr_z9 · 2026-09-02
- Markov Robotics demos sub-millimeter pick and place, highlighting zero-shot generalization for physical AGI — Scobleizer · 2026-09-02
- NVIDIA Warp hits 10M downloads; livestream to cover simulation and robotics workflows — milesmacklin · 2026-09-02