Skild's S1: robots leave the BERT era — one video prompt, zero fine-tuning

deepakpathak · x · 2026-08-26

Skild AI founder Deepak Pathak introduces S1: in-context learning for robotics via the official blog. Core thesisLanguage modeling offers a blueprint: BERT-style models needed data collection plus fine-tuning per application, while ChatGPT's pivotal leap was in-context learning — injecting new concepts via prompt without touching weights. Robotics remains stuck in the BERT era: robust new-task execution still demands tens to hundreds of hours of post-training data in deployment conditions, and research (Oh et al., 2026) shows that with sufficiently dense post-training data, from-scratch policies match post-trained foundation models — so what is pre-training for?S1 capabilities- Unseen tasks never in pre-training - 10-minute horizons - Single video prompt, no post-trainingResponding to skeptics, Pathak likens it to a natural extension of system prompts as in Claude/GPT, and notes the blog shows the robot generalizes quite far beyond the prompt. He also teases more real user cases and, eventually, home robots.

Related event: Skild AI Unveils S1, a Robot Foundation Model That Learns New Tasks From a Single Video Demo(11 posts)→

Original post →

More from Embodied

Embodied channel →