XPeng Releases Capek 0.5: An Execution-Centric VLM for Embodied Intelligence
xpeng-robotics · hf · 2026-08-10
XPeng's robotics team has released Capek 0.5, a vision-language model designed for embodied intelligence. Moving away from traditional dataset- or task-based training, it introduces an execution-centric capability taxonomy dividing embodied skills into four functional families: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification.
Architecture and Training
- RL with Verifiable Rewards (RLVR): Each capability is first trained as a dedicated specialist sharing a common backbone using RLVR.
- Consolidation: These specialists are then merged into a single inference-time model via weight-space merging and routed policy-space distillation.
Evaluation and Scale
- Instantiated at 2B and 35B-A3B scales.
- Evaluated across comprehensive benchmarks (including the new Capek-StateBench for state verification), capability retention studies, and closed-loop simulated environments.
- Capek 0.5 improves upon its initialization across most benchmarks, successfully retaining all four specialized capabilities in one checkpoint with minimal losses and transferring effectively to closed-loop tasks.
More from Embodied
- Open Source ElatoAI Runs Realtime Voice AI on ESP32 with 100+ Models — tom_doerr · 2026-08-10
- Unitree CEO Defends 219x P/E Ratio at IPO, Reveals DeepSeek Investment — 量子位 · 2026-08-10
- Rabbit R1 Review: A $199 Voice Remote for Your AI Stack, Shines When Wired to Agents — SimonBalmain · 2026-08-10
- TriWorldBench Launches as First Tri-View Evaluation for Embodied World Models — 量子位 · 2026-08-10
- Physical AI Exposes Three 'Fracture Layers'; DeepRoute Launches First Physical AI Lab in China — 量子位 · 2026-08-10
- Chinese Makers Account for 97% of Global Humanoid Robot Shipments, Report Says — rohanpaul_ai · 2026-08-10