Robotics policy keeps the vision backbone frozen and trains on a consumer GPU
mayfer · x · 2026-07-25
A robotics policy recipe that freezes a strong vision backbone, feeds all patch tokens into a small transformer policy, and uses a block-causal mask.
- No VLM or backbone fine-tuning is needed.
- The approach works with transformer-based policies such as VQ-BeT and Diffusion Policy.
- It is reported to train on a consumer GPU.
The attached diagram shows an observation trunk that encodes multi-view observations and goal inputs, then a policy head with frame-wise attention producing actions over time.
More from Embodied
- Mark Cuban says human-shaped robots may fail within 5–10 years — rohanpaul_ai · 2026-07-25
- Python visual-servoing demo shows a real closed loop on two frameworks — wightmanr · 2026-07-25
- Xpeng starts pilot production of humanoid robots as Anthropic eyes in-house chips — 创业邦 · 2026-07-25
- A prototype AI box with a keyboard points to a new mobile device form factor — dee_hw · 2026-07-25
- DroneBench shows frontier models can track people better than humans but still crash into walls — No_Call3116 · 2026-07-25
- China is unveiling humanoid AI companion robots aimed at easing loneliness — ArtificialOther · 2026-07-25