PointWAM: 3D world action model beats dexterous manipulation SOTA by 11.7 points
Chunghyun Park · hf · 2026-10-06
- PointWAM (Point World Action Model) is a 3D world action model for dexterous robotic manipulation that decomposes the world into a scene and hands, jointly forecasting both as 3D point trajectories in a shared space-time frame, then retargets predicted hand motion into robot actions.
- Unlike RGB-frame or latent world models with end-effector/joint-angle actions, this explicit, disentangled 3D representation captures spatial structure and contact geometry.
- It enables pre-training on large-scale human demonstration videos without task-specific object or keypoint selection; given a colored point cloud and a language instruction, it predicts how scene and hands co-evolve over time.
- Results: human-video pre-training lifts average DexJoCo success by 56.9 percentage points; scene-trajectory supervision adds 10.9 points over hands-only forecasting; combined, it surpasses prior SOTA on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.
More from Embodied
- $3,500 local AI PC under fire: only 24GB VRAM and priced below its own parts cost — BLUECOW009 · 2026-10-06
- Parts alone cost ~$4,500 but Ghost's Core AI PC sells for $3,499 — where's the catch? — BLUECOW009 · 2026-10-06
- Actual raw lidar image from Waymo's 6th-gen sensor suite shared online — reed · 2026-10-06
- RT-SAFE benchmark: 94.1% of 8 frontier VLMs reach goals, only 0.7% finish with zero safety events — Lianhuiq · 2026-10-06
- Vibe Robotics: Artifact Arena makes frontier models engineer competing robots — kaixhin · 2026-10-06
- JEPA-TTT cuts latent world-model prediction error 83% under dynamics shifts — JHU · 2026-10-06