AtlasVLA: Vision-Language-Action Model with Persistent Memory
CASIA-IVA-Lab · hf · 2026-08-13
A team from CASIA proposed AtlasVLA to improve the performance of Vision-Language-Action (VLA) models in embodied AI.
Replacing traditional reactive control, AtlasVLA introduces a persistent world-ego state memory mechanism to enable proactive reasoning. The architecture achieves robust long-horizon manipulation tasks using only a single wrist camera.
More from Embodied
- NVIDIA's SONIC: Scaling Laws for Natural Humanoid Whole-Body Control — yuewang314 · 2026-08-13
- Workers in India reportedly paid to film manual labor for robot training — Polymarket · 2026-08-13
- Black Forest Labs Teases FLUX 3: One Model for Image, Video, Audio, and Robotics — bennash · 2026-08-13
- Dev Integrates Vision for Local Agent Training, Explores Water Wave Computing — cephaloform · 2026-08-13
- Agility Robotics exec: Backflips easy, picking up a pen is harder — jonstephens85 · 2026-08-13
- Pi Expands Robotics Ecosystem: Hiring Generalist for Hardware & Community — minsuk_chang · 2026-08-13