RoboFollow Benchmark Exposes Instruction-Following Mirage in 9 VLA Policies
AutoLab-SJTU · hf · 2026-09-23
AutoLab at SJTU released RoboFollow, a diagnostic benchmark showing embodied agents' instruction-following ability is far weaker than success rates suggest.
- Root cause: low scene entropy — scenes admitting only one valid task make language redundant, so policies score high while barely using instructions.
- Design: high scene entropy (multiple kinematically distinct task branches per scene), a four-level (L0–L3) perturbation protocol probing spatial relations, attributes, trajectory constraints and logic, and confound-controlled Intent/Execution scoring to separate comprehension from motor execution.
- Findings: across 9 VLA and WAM policies, strong L0 performance fails to transfer to L1–L3; mitigations including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance all fail to close the gap.
- Code and dataset are open-sourced on GitHub and HuggingFace.
More from Embodied
- PixVerse world model lets you walk through AI-generated video with WASD — Aiden_Tech_Ai · 2026-09-23
- Bambu Lab launches its first laser cutter, the R1, starting at $2,499 — santoshpanda · 2026-09-23
- Porting a dual-screen device to Android 17 with Claude: 110 flashes, 3 burned resets, and eink reverse engineering — PaddleStroke · 2026-09-23
- BlackBerry QNX lands under the hood of NVIDIA's next-gen Isaac GR00T N robotics models — pdamodaran · 2026-09-23
- AI cracks 100+ open math problems but general-purpose physical AI remains missing — beffjezos · 2026-09-23
- Alibaba launches full-stack Qwen Intelligence for AI phones, Honor first to adopt — AIFlow_ML · 2026-09-23