VLA-JEPA: Training Robots to Focus on Latent Action Space
jiqizhixin · x · 2026-07-04
Research teams from USTC, Tsinghua University, and SJTU have proposed VLA-JEPA. By predicting future states in the latent space rather than at the pixel level, it avoids common issues like appearance bias and camera movement. It demonstrates better generalization and robustness than existing VLA methods on LIBERO, SimplerEnv, and real-world manipulation tasks. The arXiv paper, code, and project page are now publicly available.
More from Embodied
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11
- Musk: Cybercab certified at 165 Wh/mi, the most efficient production EV ever — elonmusk · 2026-09-11