NVIDIA's Ming-Yu Liu Explains How Physical AI Learns Across Language, Video and Action
liu_mingyu · x · 2026-09-16
NVIDIA distinguished researcher Ming-Yu Liu has released a YouTube talk explaining how Physical AI learns across three modalities: language, video, and action.
The talk focuses on training methods for embodied intelligence such as robots, and how multimodal data jointly enables physical AI to acquire real-world skills. This is a share of the official video; technical details require watching the original.
More from Embodied
- Bionic Robobird Demonstrates Nature-Mimicking Flapping-Wing Flight — TinfoilTricorn · 2026-09-16
- Robotics researcher pushes back on 'omni embodiment' hype: it's the hands, not the abstraction — chris_j_paxton · 2026-09-16
- StarVLA's VLAct trains VLAs on 16 GPUs by reshaping action representations, not data scaling — jiqizhixin · 2026-09-16
- World Labs launches Atlas: an omni world model natively spanning text, images, video and 3D — YunzhuLiYZ · 2026-09-16
- Xynova's Flex 2 bionic hand weighs 400g with 23 DOF and lifts 12kg — creatoroff · 2026-09-16
- 'GPT-6 Astra' paints Mona Lisa via SO-101 robot arm in MuJoCo demo — mishig25 · 2026-09-16