Proposal: use video generation as the high-level policy in a pi-0.7-style dual-system robot architecture
12exyz · x · 2026-08-21
Replying to a Generalist AI demo of GEN-1.5, user @12exyz proposes building a dual-system architecture similar to Physical Intelligence's pi-0.7: a language-conditioned high-level policy that produces subgoals, but swapping BAGEL for a video generation model. The original demo shows GEN-1.5 composing physical prompts from two different tasks into one continuous skill, filling in repositioning, regrasping and error-recovery motions absent from either demo. Since Generalist says its in-context learning composes video prompts and is robust to simulation and even human video, the author speculates AI-generated video could work too, unlocking language steering and longer-horizon tasks.
More from Embodied
- Lightwheel open-sources EgoSuite-Open100K: 100K hours of egocentric video for robot hands — lukas_m_ziegler · 2026-08-21
- Open-source WBC-Mjlab lands on mjswan Cloud: one policy drives G1's whole-body skills in your browser — kevin_zakka · 2026-08-21
- Humanoid robotics VC funding hits $8.7B in 2026, nearly 2x 2025's record — rohanpaul_ai · 2026-08-21
- Humanoid Robot Finger Dexterity: Bag Drops in One Second — rohanpaul_ai · 2026-08-21
- Mars is the experiment robot learning can't run on Earth — Dr_Singularity · 2026-08-21
- Galbot showcases new agile humanoid robot at WRC'26 — Distinct-Question-16 · 2026-08-21