Patch Policy: Boosting Embodied Control via Dense Visual Representations
NielsRogge · x · 2026-08-13
The post introduces a robot learning architecture named Patch Policy. Current robot policies either compress observations into a single global token or rely on heavy Vision-Language Models (VLMs), sacrificing fine-grained details and speed.\n\nPatch Policy uses a minimal architectural extension to let transformer policies consume dense pre-trained patch tokens directly. Its core is a block-causal attention mask that preserves temporal causality while processing many patch tokens.\n\nExperiments show a 40% relative improvement over policies using global-pooled representations. It also surpasses fine-tuned OpenVLA-OFT while using only about 0.7% of its parameters.
More from Embodied
- KOL Rant: 99.9% of Robotics Projects Are Just Slop — yacineMTB · 2026-08-13
- Physics-Informed Digital Twin Turns Active Particle into Autonomous Agent — bravo_abad · 2026-08-13
- Autonomous's Desk Robot Lamp Interacts with Reachy via Open-Source OS — dee_hw · 2026-08-13
- US Startups Bet on Domestic Humanoid Robots Amid China's Manufacturing Dominance — pstAsiatech · 2026-08-13
- Imagining If Robots Roamed the World of GTA — tres_pares · 2026-08-13
- Robots Now Autonomously Swapping Failed Drives in Data Centers — chris_j_paxton · 2026-08-13