INTACT by ZJU & Tsinghua: Search-Free World Model Control Directly Outputs Actions

机器之心 · wechat · 2026-07-31

Zhejiang University, Tsinghua AIR and collaborating teams proposed INTACT (INtent-To-ACTion), a search-free world model control method based on end-to-end JEPA. It directly leverages states, actions, and future results from offline trajectories to translate motion intent into a semantic interface readable by the action model.

Traditional world models typically rely on CEM or MPPI at deployment to repeatedly predict and filter from a massive number of random candidate actions, incurring significant latency. INTACT introduces a shared conditional action operator, allowing the model to directly generate an action plan without searching candidate sequences, even in complex scenarios. Small-scale search is only introduced for verification in cases like fine contact.

Experiments show that on four official LeWM tasks, INTACT achieves an average success rate of 95.33% with zero search after training for only 1 epoch. With optional small-scale local verification, the success rate further improves to 96.86%. Compared to the traditional 9,000 candidate sequences, sampling volume is reduced by 23.44x, and planning latency is reduced to 1/300th.

Related event: INTACT: A Search-Free World Model by ZJU and Tsinghua(2 posts)→

Original post →

More from Embodied

Embodied channel →