ACG-Bench Probes Dual-Arm VLA Generalization; AE-VLA Lifts Success from 3% to 21.5%
Zaibin Zhang · hf · 2026-10-06
ACG-Bench studies arm-wise compositional generalization in dual-arm vision-language-action models—recomposing familiar atomic skills in new ways across arms.
- Benchmark: 23 task-condition pairs across 8 task families, with 6 in-domain conditions and 17 unseen compositions (reordering, synchronization, their combination, cross-task composition), requiring task goals plus physical milestones and order/timing constraints.
- Architecture study: Using π{0.5} as backbone, combining arm-token grouping, SkillLoRA adapters, and arm-wise attention yields AE-VLA at 21.53% simulation generalization success vs 2.94% for single π{0.5} and 5.53% for two independently controlled policies.
- Physical robots: On SO101 arms, AE-VLA reaches 39.00% mean success over five unseen conditions vs 10.00% for the strongest baseline.
More from Embodied
- The robotics data business's biggest bottleneck is operational, not technical — paigeinsf · 2026-10-06
- GPT-2-powered robot Samson explores, gets a little confused but has fun — MikePFrank · 2026-10-06
- VC: almost none of real robot deployment data flows back into training — carrycooldude · 2026-10-06
- RealtimeWAM: One-Step Asynchronous World Action Model Delivers 25x Speedup with <1% Accuracy Loss — NanyangTechnologicalUniversity · 2026-10-06
- ReSteer open-sourced: fixing VLA policies that ignore mid-execution instruction switches — siddkaramcheti · 2026-10-06
- DIY robotic backboard uses computer vision to make every beginner shot go in — TinfoilTricorn · 2026-10-06