Can interleaved video and text pretraining merge world models for acting and talking?
voooooogel · x · 2026-08-20
Explores a hypothetical training method involving interleaved video and text pretraining with separate output heads. It speculates whether this approach causes world models to merge, and whether subsequent cheap interleaved fine-tuning, inspired by Pantograph's action interleaving, could yield a model capable of both conversation and action.
More from Research
- Weaviate Podcast Discusses Reranking in Deeper Pools — CShorten30 · 2026-08-20
- NLP Classic 'Speech and Language Processing' Releases August 2026 Update — srchvrs · 2026-08-20
- New Paper Formulates Global Workspace Theory Using Control-Theoretic Math — Jack_W_Lindsey · 2026-08-20
- Meituan's search 3.0: LLM semantic embeddings lift long-tail NDCG by 2.21pp across three iterations — 美团技术团队 · 2026-08-20
- SkillForge: Self-Distilling Agents for Project-Specific Bug Fixing — SJTU · 2026-08-20
- FM-Bench Evaluates Long-Horizon Agent Management via Football Club Simulation — Tianyou Wang · 2026-08-20