TEMPO adds temporal context to VLAs, lifting robot bottle handover success from 44% to 74%
_krishna_murthy · x · 2026-10-08
UC Irvine researchers present TEMPO (CoRL 2026 oral), fixing why VLA models fail at dynamic manipulation like catching balls or pouring into a moving pot.
- Diagnosis: single-frame observations cause motion ambiguity (can't predict moving objects) and state aliasing (visually similar states needing different actions); the bottleneck is missing temporal context, not model scale or latency.
- Method: augment a frozen pretrained VLA with a motion summary from a frozen video foundation model plus compact proprioceptive history—no backbone changes, minimal overhead.
- Results: 1,500+ physical evaluations (50 runs per success rate); Bottle Handover improves from 44% to 74%, the only method to solve state aliasing.
- Open: paper, code, data, and TEMPO-Bench (50k+ annotated frames) all released.
More from Embodied
- Berkeley researcher launches RPG: guided self-improvement for embodied agents — berkeley_ai · 2026-10-08
- Surface Laptop Ultra ships Oct 16 at $2,599 with magnetic USB-C 'Magnetic Connect' charging — tomwarren · 2026-10-08
- Surface Laptop Ultra launches Oct 16 starting at $2,599 with Nvidia RTX Spark — tomwarren · 2026-10-08
- Atomic Machines builds an AI 'matter compiler' to make index-finger-sized circuit breakers — ZoubinGhahrama1 · 2026-10-08
- Jaguar Type 01 debuts on NVIDIA Hyperion platform, software passed 150,000 tests — nvidia · 2026-10-08
- ProHand 2 Power dexterous robot hand debuts at IROS 2026 — imankitgoyal · 2026-10-08