VLA generalization may fail because instructions are too short, researcher argues

YouJiacheng · x · 2026-10-11

Citing a talk at Alignment 2026, YouJiacheng argues a key reason VLAs generalize poorly is data presentation: models map short text instructions to long, complex trajectories, so instructions lack the information to explain the trajectories and the model overfits to environment details instead. He suggests much richer instructions describing how to actually perform the task, drawing a parallel to image/video generation where core models rely on increasingly detailed prompts — with a model handling the short-request-to-detailed-prompt expansion.

Original post →

More from Embodied

Embodied channel →