Can interleaved video and text pretraining merge world models for acting and talking?

voooooogel · x · 2026-08-20

Explores a hypothetical training method involving interleaved video and text pretraining with separate output heads. It speculates whether this approach causes world models to merge, and whether subsequent cheap interleaved fine-tuning, inspired by Pantograph's action interleaving, could yield a model capable of both conversation and action.

Original post →

More from Research

Research channel →