Tsinghua's ST-WAM Boosts Robot Manipulation Robustness Under Visual Shifts
Tsinghua · hf · 2026-08-05
This paper proposes the Semantic-Temporal World Action Model (ST-WAM) to tackle the lack of robustness in World Action Models (WAMs) under visual distribution shifts.
- Problem Addressed: Overcomes "Training-Distribution Hallucination," where models hallucinate training-domain content instead of adhering to the current scene.
- Architecture: Uses DINOv3 as a shared semantic representation for future prediction and history retrieval, featuring Dual-Space Future Experts (DSFE) and Current-Anchored Intent Retrieval (CAIR).
- Performance: Achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0. Improves zero-shot real-world success rates under visual shifts from 25.8% to 61.5%.
More from Embodied
- Google Launches Gemini Robotics ER 2 for Real-Time Multi-Robot Control — heypearlai · 2026-08-05
- Portable 64-Channel Millimeter-Wave Camera with EO Conversion Unveiled — jwt0625 · 2026-08-05
- Robots Find Exact Stop Points Better Than Overall Progress: Gemini Eval — Crescitaly · 2026-08-05
- Unitree G1 Robot Gets Unified AI Brain Upgrade for Enhanced Perception and Movement — Scobleizer · 2026-08-05
- Scoble: US Robots Risk Losing Global Market if They Lag Behind China — Scobleizer · 2026-08-05
- GE Appliances Integrates Texas Instruments Chips for Next-Gen AI-Embedded Appliances — emmanuelvivier · 2026-08-05