NVIDIA's Ming-Yu Liu on Cosmos 3: one model for video, understanding and robot actions

Machine Learning Street Talk · youtube · 2026-09-16

Machine Learning Street Talk interviews Ming-Yu Liu, who leads NVIDIA's Cosmos research. The left-turning car that opens the episode was never filmed — Cosmos 3 generated it. Liu explains how one model can describe video, generate it, and produce robot actions.

Architecture: a vision language model reasons one token at a time; its weights initialize a bidirectional diffusion generator for video, audio and action, with a shared temporal position scheme aligning signals running at different rates.

World models: Liu treats them as a toolkit — forward dynamics, inverse dynamics and policy trained together under a capacity limit so each helps the others.

Data transfer: plentiful first-person human video carries over to data-starved robots; a Cosmos model post-trained on DROID is a strong starting point for pick-and-place policies.

Testing (the most practical thread): a neural simulator doesn't need accurate success rates — it only needs to rank policies like the real world would, narrowing checkpoints for real trials. Cosmos Dreams applies this closed-loop idea to driving and robotics; Liu argues safety matters even more for humanoids around children and pets.

Sizes: Super, Nano and Edge (Edge targets Jetson Thor, Orin, DGX Spark), with open weights, code and data. Paid partnership with NVIDIA.

Related event: NVIDIA Cosmos 3: One World Model for Video Generation, Understanding and Robot Action(3 posts)→

Original post →

More from Embodied

Embodied channel →