Anchor-Align: Recovering OOD Generalization in VLA Fine-tuning
_krishna_murthy · x · 2026-08-26
Researchers from UIUC and others introduce Anchor-Align, a method to address the degradation of generalization capabilities in Vision-Language-Action (VLA) models after Behavior Cloning (BC) fine-tuning.
Core Problem: Fine-tuning VLMs with BC progressively overwrites pretrained representations essential for visual and semantic generalization. Simple co-training on web image-text data fails to resolve language-action misalignment.
Solution: Anchor-Align augments BC with two objectives:
- Vision-Language Anchoring: Distills layer-wise representations from a frozen VLM copy to prevent drift.
- Language-Action Alignment: Converts action targets into discrete motion-direction labels and jointly trains language and action prediction on the same observation.
Results: On a physical xArm7 robot, Anchor-Align improves success rates from 28% → 54% and 37% → 60% across two VLA architectures. Consistent improvements in OOD perturbations, perceptual robustness, and long-horizon tasks were demonstrated in simulations (LIBERO, CALVIN).
More from Embodied
- OOMWOO Open-Source Robot Vacuum: Raspberry Pi, ROS 2, and Fully Local Control — JeremyCMorgan · 2026-08-26
- Mechanical Prosthetic Hand Designed for Amniotic Band Syndrome — _Stocko_ · 2026-08-26
- Robot Racing: More Entertaining Than Car Races — cixliv · 2026-08-26
- Robotics training debate: sim distillation vs. human behavior cloning — KyleMorgenstein · 2026-08-26
- Researchers discuss robot world models at IAS workshop in Darmstadt — breadli428 · 2026-08-26
- Chinese humanoid robot sales dwarf US counterparts — teortaxesTex · 2026-08-26