CSWAM paper adds a V-JEPA 2.1 causal semantic expert to boost world action model OOD generalization
zhenjun_zhao · x · 2026-10-01
The CSWAM (Causal Semantic World Action Model) paper targets FastWAM-style world action models' poor generalization under visual distribution shifts:
- Problem: action-only inference is efficient but reconstruction-oriented representations overfit to appearance details, and the lack of observation history makes it hard to identify task-relevant state changes in unfamiliar conditions.
- Method: augment FastWAM with a causal semantic expert built on V-JEPA 2.1, which provides temporally grounded representations of semantic state changes and motion. The expert learns future evolution from a sparse history of observations and shares the history-derived context with both video and action streams via causal attention.
- Inference: action denoising is conditioned on the current video state and observed semantic history while keeping efficient action-only inference.
- Results: simulated and real-robot experiments show improved generalization under distribution shifts, especially after embodied pretraining.
More from Embodied
- Robot demo shows tactile sensing across fingers and entire palm — CyberRobooo · 2026-10-01
- LATENT wins IROS 2026 award: humanoid robots rally at human level from imperfect motion data — chris_j_paxton · 2026-10-01
- EngineAI T800 hailed as China's most dynamic robot, but fell more than any other at IROS — chris_j_paxton · 2026-10-01
- Custom Retro Hardware Is a Whole New World with AI for Hardware Design — pvncher · 2026-10-01
- Researcher dismantles Figure's IP excuse for destroying F.02: "awful marketing" — MarwaEldiwiny · 2026-10-01
- Autonomous's $20,900 Solar WorkPod Offers a Backyard Office with Dual RTX 5090 AI Datacenter — dee_hw · 2026-10-01