INTACT: Search-Free World Models via Isomorphic Intent-to-Action
zju · hf · 2026-07-31
Forward latent world models predict how actions change a scene but recover desired actions only through expensive test-time search. This work introduces INTACT (INtent-To-ACtion), an end-to-end JEPA architecture turning action-labeled trajectories into a deployable intent-to-action interface.
- Architecture: Maintains isomorphism between local and goal motion-intent input graphs through an identical four-slot grammar and shared parameters. Supports intact transfer from RGB evidence to action-effective latent intent coordinates.
- Search-Free Policy: Uses the conditional mean as a robust search-free policy without pointwise latent matching. Optional local CEM reduces sampling by 23.44x.
- Performance: On four official LeWM tasks, one-epoch, zero-search models reach 85.78% to 100% success. Using 384 instead of 9,000 candidate sequences achieves 96.86% macro success, improving pure CEM by 16.00 points.
Related event: INTACT: A Search-Free World Model by ZJU and Tsinghua(2 posts)→
More from Embodied
- New Review on Opportunities for Legged Robots by Jonas Frey et al. — ChongZzZhang · 2026-07-31
- Qlayers Robots Replace Hazardous Manual Labor, 6x Faster Tank Coating — lukas_m_ziegler · 2026-07-31
- 50+ Institutions Unite to Build Open-X-Tactile, World's Largest Tactile Dataset — RemiCadene · 2026-07-31
- Qualcomm Pushes for AI Gaming Era: On-device Compute Shrinks Dev Time to 5 Days — 智东西 · 2026-07-31
- Harness VLA: Boosting Frozen Robot Models Without Retraining — jiqizhixin · 2026-07-31
- Expert Reveals the Golden Rule of Robotics Demos: What You See Is What You Get — chris_j_paxton · 2026-07-31