Video2Skill benchmark: most of 19 open-source VLMs fail at streaming embodied skill discovery
Jianshu Zhang · hf · 2026-10-07
The paper formulates Streaming Embodied Skill Discovery (SESD) and introduces the Video2Skill benchmark: models watch a video stream and maintain a persistent skill library that shapes later decisions.
- Covers robot tabletop manipulation and human kitchen activity, testing event localization, grouping similar transformations, and reuse-vs-create decisions
- Across 19 open-source VLMs, many group events at near-chance level; scale doesn't consistently help
- Errors depend on coupling: joint models merge distinct transformations into one skill, while text-based library updates duplicate recurring ones
- Supervised fine-tuning with the authors' CLaRe rebalancing improves grouping but hits a deeper bottleneck: libraries stall below half the reference size, and unseen transformations are localized but almost never given a new skill
The central challenge: recognizing when existing skills are insufficient.
More from Embodied
- Tesla's Cybercab rests on aggressive FMVSS interpretation NHTSA could reject — binarybits · 2026-10-07
- Waymo Colors Inside FMVSS Lines; Tesla Bets on Regulatory Favor, Analyst Argues — binarybits · 2026-10-07
- 19-Year-Old Founder Zain Raises $11M Seed Led by a16z to Ship Personal AI Computer Core — nick_linck · 2026-10-07
- Father-son team launches Rhem, a companion robot helping aging parents call, book and track health — ycombinator · 2026-10-07
- Sentdex finds decision language models fall short in real robotics pipelines vs RL/VLAs — Sentdex · 2026-10-07
- NeurIPS paper adds probabilistic uncertainty quantification to robot memory for better retrieval — lucacarlone1 · 2026-10-07