NUS Show-Harness: A Semantic Interface Lets VLMs Control Real Robots Zero-Shot
机器之心 · wechat · 2026-09-27
NUS ShowLab's Show-Harness skips training bigger robot policies and instead inserts an embodied harness — a semantic action interface — between foundation models and robots: VLMs reason over readable units (move forward/left, rotate, grasp, release) while per-robot interpreters map them to bounded low-level commands, enabling cross-embodiment by swapping interpreters only. Decisions run in a perceive–reason–act loop rather than blind long-horizon plans.
Real-robot results: zero-shot frontier VLMs hit 89% (cross-task), 100% (cross-environment), 93% (cross-embodiment, Franka + AgileX incl. dual-arm) vs. baselines at 57–65%. A lightly fine-tuned Qwen3.5-2B (rank-64 LoRA, 3% params) transfers 230 sim demos to 13/20 real Franka tasks where trainable VLA baselines scored 0/20. An ablation shows semantic names alone keep success at 90–95%, but removing both names and physical conventions crashes it to 5% — stable symbol-to-effect grounding is the key.
The companion GUMI GUI renders the same action units for humans (keyboard teleop) and Computer-Use agents alike, while logging (observation, action) data for training. Paper, code, models, and data are open-sourced, plus a survey on multimodal embodied agents.
More from Embodied
- Robotics builder: the hard part is finding use cases someone will actually pay for — TajyMany · 2026-09-27
- Modern Robotics: the graduate textbook on mechanics, planning and control — blaizedsouza · 2026-09-27
- Dev: no point getting LASIK in 2026 — AI glasses will bring face-worn superpowers — willcb · 2026-09-27
- Audio codec models make surprisingly good action tokenizers for dexterous robot arms — chris_j_paxton · 2026-09-27
- Flying cars finally arrive: rider shares first-hand eVTOL takeoff video — chris_j_paxton · 2026-09-27
- Opus 5.5 demos nail spatial perception, but tactile sensing remains VLA robotics' frontier — bookwormengr · 2026-09-27