Attacking the First Denoising Step is Enough to Break Flow-Matching VLAs

Hoseong Tae · hf · 2026-08-06

Researchers introduced DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch attack targeting flow-matching Vision-Language-Action (VLA) models like pi0.

The paper reveals that the perceived robustness of these models against adversarial perturbations is largely illusory. Counterintuitively, attacking only the very first denoising step is both stronger and cheaper than attacking a wider window of steps. Across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks for pi0 and pi0.5 using just a small single patch, significantly outperforming action- and embedding-space attack baselines.

Original post →

More from Embodied

Embodied channel →