Proposal: use video generation as the high-level policy in a pi-0.7-style dual-system robot architecture

12exyz · x · 2026-08-21

Replying to a Generalist AI demo of GEN-1.5, user @12exyz proposes building a dual-system architecture similar to Physical Intelligence's pi-0.7: a language-conditioned high-level policy that produces subgoals, but swapping BAGEL for a video generation model. The original demo shows GEN-1.5 composing physical prompts from two different tasks into one continuous skill, filling in repositioning, regrasping and error-recovery motions absent from either demo. Since Generalist says its in-context learning composes video prompts and is robust to simulation and even human video, the author speculates AI-generated video could work too, unlocking language steering and longer-horizon tasks.

Original post →

More from Embodied

Embodied channel →