Fine-Grained Annotation Boosts VLA Performance

JeffDean · x · 2026-07-12

This repost introduces a VLA model study combining vision + proprioception: by using finer-grained sub-task annotations, the model achieved a new SOTA in sub-task generation and generalized better across different embodiments.

The post highlights two specific results: achieving 93.1 F1@50 on REASSEMBLE and 98.6 on the Amazon Robotics blade insertion task. The original thread is linked for further reading.

Original post →

More from Embodied

Embodied channel →