SimpleMemVLA feeds full video history to a VLM, beating dedicated memory modules for long-horizon robot manipulation

openbmb · hf · 2026-09-09

openbmb released SimpleMemVLA, a simple native-video memory approach for vision-language-action models. It feeds intact timestamped video history directly into a pretrained VLM backbone and uses hidden states to inform a flow-matching action head, enabling long-horizon manipulation that outperforms dedicated memory modules.

Original post →

More from Embodied

Embodied channel →