Berkeley and DeepMind Propose Instruct-to-Act: Making World-Model Controllers Language-Instructable

berkeley_ai · x · 2026-09-22

Zineng Tang et al. (UC Berkeley, Vector Institute, Google DeepMind) present Instruct-to-Act at COLM 2026, a system that decouples planning and control for language-instructable agents.

Key ideas:

The paper reports results across embodied environments, scaling trends for larger controllers and VLMs, and ablations on instruction cadence, planning frequency, and online vs. offline planning latency. Code and paper are available.

Original post →

More from Embodied

Embodied channel →