Barron talk Part 3: in the limit, robotics probably shouldn't use explicit 3D — just predict actions

jon_barron · x · 2026-08-31

Part 3 of Jon Barron's talk makes a pointed claim: in the limit, robotics probably shouldn't use explicitly 3D representations at all, and should instead focus on the easier task of directly predicting actions. This echoes the "Bitter Lessons" thesis — rather than building structured intermediate representations, let models learn outputs end-to-end, aligning with the current VLA (vision-language-action) trend.

Related event: Jon Barron's CVPR talk: when explicit 3D representations are worth it(10 posts)→

Original post →

More from Embodied

Embodied channel →