BridgeVLA++ Boosts 3D Robotic Manipulation with Spatio-Temporal Memory Architecture
Peiyan Li · hf · 2026-08-06
Researchers introduced BridgeVLA++, a memory-augmented Vision-Language-Action (VLA) framework for 3D robotic manipulation. It addresses limitations in existing 3D VLA models, which are typically data-hungry, struggle with out-of-distribution generalization, and lack explicit memory of past observations.
Building on its predecessor's method of projecting point clouds into multi-view images, BridgeVLA++ integrates a unified spatio-temporal memory architecture that models both persistent spatial context and temporal interaction history. Experiments show the framework achieves strong performance on spatial tasks and sets a new state-of-the-art on two challenging memory-dependent benchmarks. It also proves effective in bimanual manipulation settings and on real-world robotic platforms, demonstrating robust scalability across tasks and hardware.
More from Embodied
- Indie Dev Uses AI to Design Circuit Boards, Aims to Crack Production and Sell Hardware — pramodk73 · 2026-08-06
- US Robot Ban Hits Startups: Requires Over 65% Domestic BOM Sourcing — mattfreed · 2026-08-06
- Nori L3 Dual-Arm Home Robot Launched at $1,688 — CyberRobooo · 2026-08-06
- PowerBot's Multi-Sport AI Coach Robot Raises Over $4M on Kickstarter — 创业邦 · 2026-08-06
- Beijing Deploys 72 Park Robots for Patrolling, Cleaning, and Pest Control — pstAsiatech · 2026-08-06
- Amazon's Zoox to Launch Paid Fully Driverless Robotaxi Rides in Las Vegas — Polymarket · 2026-08-06