VLA-JEPA: Training Robots to Focus on Latent Action Space
jiqizhixin · x · 2026-07-04
Research teams from USTC, Tsinghua University, and SJTU have proposed VLA-JEPA. By predicting future states in the latent space rather than at the pixel level, it avoids common issues like appearance bias and camera movement. It demonstrates better generalization and robustness than existing VLA methods on LIBERO, SimplerEnv, and real-world manipulation tasks. The arXiv paper, code, and project page are now publicly available.
More from Embodied
- Teachers decry plan to put a humanoid robot in a New York high school — nordicinst · 2026-07-27
- NUS builds a soft force sensor that drives actuators without electronics or power — CurieuxExplorer · 2026-07-27
- Chelsea Finn says robot RL is bottlenecked by physical rollout cost, not algorithms — ycombinator · 2026-07-27
- Robot goes to the fridge and fetches a beer — Darpinian · 2026-07-27
- Researchers show digital circuits can be replicated with knitted fabric — mtizard · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27