StarVLA's VLAct trains VLAs on 16 GPUs by reshaping action representations, not data scaling
jiqizhixin · x · 2026-09-16
StarVLA's VLAct (Beyond Data Scaling: Representation-Centric Continued Pre-training for VLAs) argues VLA progress can come from better representations instead of more data — crucial since robot data can't be scraped and every trajectory requires real hardware and teleoperation.
- Continued pre-training uses only 16 GPUs on open-source data.
- Key technique: an Orthogonal Finetuning (OFT) head that directly shapes the backbone's action representations.
- Evidence matters: an OFT head improves from 61.7% to 75.8%, but swapping the same backbone to a PI head loses the gains — the improvement lives in the representation, not the weights alone.
More from Embodied
- Travis Kalanick: Tesla is 'the Google of this era' in the physical AI age — rohanpaul_ai · 2026-09-16
- Slovenia's Deputy PM tries Tesla FSD on public roads: 'doesn't get tired, doesn't fall asleep' — elonmusk · 2026-09-16
- DRS-VPT: feed-forward camera pose estimation from a point cloud scan and a single image — kwangmoo_yi · 2026-09-16
- Bionic Robobird Demonstrates Nature-Mimicking Flapping-Wing Flight — TinfoilTricorn · 2026-09-16
- Robotics researcher pushes back on 'omni embodiment' hype: it's the hands, not the abstraction — chris_j_paxton · 2026-09-16
- World Labs launches Atlas: an omni world model natively spanning text, images, video and 3D — YunzhuLiYZ · 2026-09-16