SimpleMemVLA feeds full video history to a VLM, beating dedicated memory modules for long-horizon robot manipulation
openbmb · hf · 2026-09-09
openbmb released SimpleMemVLA, a simple native-video memory approach for vision-language-action models. It feeds intact timestamped video history directly into a pretrained VLM backbone and uses hidden states to inform a flow-matching action head, enabling long-horizon manipulation that outperforms dedicated memory modules.
More from Embodied
- CoRL 2026 to host 'Pretrain to Adapt' workshop: strongest robot policies aren't easiest to adapt — canondetortugas · 2026-09-09
- Enactic founder exits humanoid robotics startup after OpenArm open-source success — hiro_yams · 2026-09-09
- Apple acquires Sonera, betting on magnetic-field sensing to read brain activity in wearables — asymco · 2026-09-09
- You can now control a stepper motor for just $1.20 — _Stocko_ · 2026-09-09
- World Labs demos real-time next-view prediction with accurate 3D consistency — drfeifei · 2026-09-09
- 18-year robotics veteran: the gap between demo intelligence and field intelligence never closed — yongqianme · 2026-09-09