NVIDIA's Ming-Yu Liu Explains How Physical AI Learns Across Language, Video and Action

liu_mingyu · x · 2026-09-16

NVIDIA distinguished researcher Ming-Yu Liu has released a YouTube talk explaining how Physical AI learns across three modalities: language, video, and action.

The talk focuses on training methods for embodied intelligence such as robots, and how multimodal data jointly enables physical AI to acquire real-world skills. This is a share of the official video; technical details require watching the original.

Original post →

More from Embodied

Embodied channel →