NVIDIA's Ming-Yu Liu Explains Cosmos 3: One World Model for Video, Audio and Robot Actions
Machine Learning Street Talk · rss · 2026-09-16
Machine Learning Street Talk interviews Ming-Yu Liu, who leads NVIDIA's Cosmos research, on Cosmos 3—the omnimodal world model that generated the never-filmed left-turning car opening the episode.
Highlights:
- Architecture: a vision language model reasons token by token; its weights initialize a bidirectional diffusion generator producing video, audio and robot actions, aligned by a shared temporal position scheme.
- World models as tools: forward dynamics, inverse dynamics and policy trained together under a capacity limit so each helps the others.
- Data: plentiful first-person human video transfers to data-starved robots; Cosmos post-trained on DROID is a solid starting point for pick-and-place policies.
- Testing: a neural simulator only needs to rank policies like the real world would, letting teams pick checkpoints for real trials; Cosmos Dreams applies this closed-loop idea to driving and robotics.
- Three sizes—Super, Nano, Edge (Edge targets Jetson Thor, Orin, DGX Spark)—with open weights, code and data. Episode is an NVIDIA paid partnership.
More from Embodied
- Strapping an LLM onto the Jev robot to turn vague instructions into precise ones — daniel_mac8 · 2026-09-16
- Maker Arm ships first batch: open-source 3D-printed 6-axis robot arm from $999 — IsaiahBallah · 2026-09-16
- Chinese company unveils lifelike robotic fish that could blend into real fish schools for underwater surveillance — Polymarket · 2026-09-16
- Dual ALOHA robot demos unlock interlocked parts and thread a rope through three rings — jiqizhixin · 2026-09-16
- Force Origin DM0.5 tops six embodied-AI benchmarks, sweeps all four RoboColiseum boards — 机器之心 · 2026-09-16
- Robotics Researcher: The GPT of Robotics Will Just Be GPT — chris_j_paxton · 2026-09-16