ActionPiece rethinks action tokenization for VLA models, hits 94.8% on LIBERO
DeepCybo · hf · 2026-09-17
The paper argues that pointwise reconstruction metrics like MSE fail to capture whether action tokenizers in autoregressive VLA models preserve context-dependent adjustments after compression—similar actions may collapse while needed adjustments are diminished or reversed.
The authors introduce physical rank consistency (PRC) to measure how well local physical distance rankings survive reconstruction, and propose ActionPiece, which preserves physical action relationships via joint supervision of representation learning and quantization: rank preservation constrains near-far ordering in encoder/quantized feature distances, while quantization regularization applies the same ordering to codeword assignments.
Under the same Qwen3-VL-4B policy setup, ActionPiece reaches 94.8% on LIBERO, 68.8% on unseen LIBERO-Plus, 71.9% on SimplerEnv, and 51.5% on VLA-Arena L0-L2, with ablations showing both objectives jointly improve PRC and policy success.
More from Embodied
- Tesla Optimus passes CAPTCHA designed to prove you're not a robot — TansuYegen · 2026-09-17
- Bionic robotic fish flexes body like a real fish at China trade fair — rohanpaul_ai · 2026-09-17
- GPT-6 Astra reportedly pretrained on 100k+ GPUs at Stargate, with big real-to-sim implications — erwincoumans · 2026-09-17
- UBTECH launches hyper-realistic U1 companion humanoid robots, sparking privacy questions — TansuYegen · 2026-09-17
- The egocentric data boom: Maxinsights has delivered 2M+ hours for robot training — 机器之心 · 2026-09-17
- Ulysses co-founder on Mako modular underwater robots and seabed restoration — Scobleizer · 2026-09-17