ActionPiece rethinks action tokenization for VLA models, hits 94.8% on LIBERO

DeepCybo · hf · 2026-09-17

The paper argues that pointwise reconstruction metrics like MSE fail to capture whether action tokenizers in autoregressive VLA models preserve context-dependent adjustments after compression—similar actions may collapse while needed adjustments are diminished or reversed.

The authors introduce physical rank consistency (PRC) to measure how well local physical distance rankings survive reconstruction, and propose ActionPiece, which preserves physical action relationships via joint supervision of representation learning and quantization: rank preservation constrains near-far ordering in encoder/quantized feature distances, while quantization regularization applies the same ordering to codeword assignments.

Under the same Qwen3-VL-4B policy setup, ActionPiece reaches 94.8% on LIBERO, 68.8% on unseen LIBERO-Plus, 71.9% on SimplerEnv, and 51.5% on VLA-Arena L0-L2, with ablations showing both objectives jointly improve PRC and policy success.

Original post →

More from Embodied

Embodied channel →