FailBank Turns Runtime Shield Feedback into Persistent VLA Policy Gains (+25.4 Points Success)

notredame · hf · 2026-10-05

FailBank is a four-stage self-evolution framework that converts runtime feedback into persistent policy improvement for vision-language-action models, addressing the persistent policy-shield mismatch that pure runtime shields create.

On the VLA-Arena benchmark across two VLA backbones: +8.5 and +6.9 points task success over base policies while cutting policy-induced cumulative cost by 35.6% and 23.8%; versus runtime shielding, +25.4 and +9.5 points at comparable cost. Runtime feedback can serve as persistent policy supervision rather than a temporary action constraint.

Original post →

More from Embodied

Embodied channel →