Patch Policy beats a fine-tuned 7B VLA by 18% with 0.7% of the parameters
ylecun · x · 2026-07-23
Researchers introduce Patch Policy, a minimal architectural extension for robot policies that lets transformer-based agents consume dense ViT patch tokens directly instead of collapsing vision into a single vector first.
- The paper argues that pretrained ViTs preserve rich spatial detail that standard policies throw away.
- Patch Policy reaches 18% better performance than a fine-tuned 7B VLA while using only about 0.7% of its parameters.
- The result is presented as a path to more robust and precise manipulation without needing a billion-parameter vision-language model.
More from Embodied
- NVIDIA says physical AI needs exabyte-scale simulation data at SIGGRAPH 2026 — jonstephens85 · 2026-07-23
- Musk says Optimus aims to be the first humanoid robot useful in daily life — XFreeze · 2026-07-23
- NVIDIA says its open-source robotics simulator cut training from 5 hours to under 2 minutes — imjustnewatai · 2026-07-23
- Another repeat of the Vicarious vs. da Vinci robotics success story — AIandDesign · 2026-07-23
- Tesla says Optimus production lines are being installed, with robot output due soon — Polymarket · 2026-07-23
- Tesla says Cybercab production has started, Robotaxi is live in 7 US metros — yunta_tsai · 2026-07-23