BridgeVLA++ Boosts 3D Robotic Manipulation with Spatio-Temporal Memory Architecture

Peiyan Li · hf · 2026-08-06

Researchers introduced BridgeVLA++, a memory-augmented Vision-Language-Action (VLA) framework for 3D robotic manipulation. It addresses limitations in existing 3D VLA models, which are typically data-hungry, struggle with out-of-distribution generalization, and lack explicit memory of past observations.

Building on its predecessor's method of projecting point clouds into multi-view images, BridgeVLA++ integrates a unified spatio-temporal memory architecture that models both persistent spatial context and temporal interaction history. Experiments show the framework achieves strong performance on spatial tasks and sets a new state-of-the-art on two challenging memory-dependent benchmarks. It also proves effective in bimanual manipulation settings and on real-world robotic platforms, demonstrating robust scalability across tasks and hardware.

Original post →

More from Embodied

Embodied channel →