Steering VLA robot models at inference time lifts grasp rate from 14% to 74% without retraining
drmapavone · x · 2026-09-23
A new paper asks whether robot foundation models (VLAs) can be steered without retraining.
- Drawing on control-theoretic notions of observability and controllability, the authors show state- and action-relevant information can often be linearly predicted from VLA internal representations
- These representations can be steered at inference time to change robot behavior, with no fine-tuning
- On a real robot, the approach raised the preferred handle-grasp rate from 14% to 74% with only 1% inference overhead
- The broader idea: representation-level control as a lightweight interface for adapting robot foundation models to new preferences and constraints
Joint work with Hugo Buurmeijer, Carmen Amo Alonso and Swann Aiden; paper is publicly available.
More from Embodied
- eidon-ai releases tracker-pov robotics video dataset, 10K-100K entries — eidon-ai · 2026-09-23
- Dev turns a 90s TV into an AI TV with Raspberry Pi, mic, camera and GPT Realtime — OpenAIDevs · 2026-09-23
- Robot arm tests rank GPT-6 variants: better reasoning means better manipulation at 30x the cost — YuXiang_IRVL · 2026-09-23
- Frontier AI models can control robots to follow harmful requests, NBC News reports — ycombinator · 2026-09-23
- 'Robot-use' works on humanoids: coding agents zero-shot control a full humanoid for pick-and-place — ai · 2026-09-23
- REAL: an embodied agent that explores real rooms and asks clarifying questions, accepted at ECCV 2026 — jiqizhixin · 2026-09-23