Attacking the First Denoising Step is Enough to Break Flow-Matching VLAs
Hoseong Tae · hf · 2026-08-06
Researchers introduced DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch attack targeting flow-matching Vision-Language-Action (VLA) models like pi0.
The paper reveals that the perceived robustness of these models against adversarial perturbations is largely illusory. Counterintuitively, attacking only the very first denoising step is both stronger and cheaper than attacking a wider window of steps. Across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks for pi0 and pi0.5 using just a small single patch, significantly outperforming action- and embedding-space attack baselines.
More from Embodied
- Embodied AI Alliance: Riemann Dynamics & Partners Build 1M-Hour Data Foundation — 量子位 · 2026-08-06
- Hands-on with UBTech Walker S2: Low Payload, Inflexible Hands, Weird Gait — lukas_m_ziegler · 2026-08-06
- Keychron Teases Tiny Wireless AI Companion Hardware for Vibe Coding — emax · 2026-08-06
- Beihang University Unveils 2cm Micro-Bot Bug That Moves Fast and Senses Sound — lukas_m_ziegler · 2026-08-06
- Dev Uses Claude Opus to Write C Code Driving ESP32 S3 Hardware — petewoodbridge · 2026-08-06
- NeurIPS 2026 Call for Papers: 2nd Embodied Spatial Reasoning Workshop — hunarbatra · 2026-08-06