AtlasVLA: Vision-Language-Action Model with Persistent Memory

CASIA-IVA-Lab · hf · 2026-08-13

A team from CASIA proposed AtlasVLA to improve the performance of Vision-Language-Action (VLA) models in embodied AI.

Replacing traditional reactive control, AtlasVLA introduces a persistent world-ego state memory mechanism to enable proactive reasoning. The architecture achieves robust long-horizon manipulation tasks using only a single wrist camera.

Original post →

More from Embodied

Embodied channel →