τ(0)-VLA uses 40,115 hours of robot data to tackle long-horizon tasks
机器之心 · wechat · 2026-07-27
A long WeChat article breaks down τ(0)-VLA, a new embodied AI model aimed at long-horizon robotic tasks. The key idea is to split planning and execution: a high-level “slow thinking” policy handles task decomposition and decision making, while a low-level “fast execution” policy handles real-time control.
What the model adds
- Introduces test-time computation and a world-model-guided planning loop for embodied decision making
- Uses a four-part high-level stack: proposal model, world model, value model, and reflection model
- Builds a beam-search-like mechanism over sub-tasks to reason about chain effects in the physical world
Training and scale
- Trained on 40,115 hours of real robot interaction data
- Includes 20,000+ hours of on-robot data across multiple robot forms and public datasets
- Supports multiple embodiments via a unified 40-dimensional action space and masking for different robots
Reported results
- On AGIBOT G1, hierarchical planning raises average success rate from 27.5% to 45.0% and progress from 80.10% to 87.85%
- The model also shows strong direct-execution performance across several manipulation tasks, with some tasks reaching 10/10 success
The article argues that embodied AI is moving from short demo actions toward long-horizon real-world completion, and that “thinking before acting” may be the key bottleneck to solve.
Related event: τ₀-VLA Tackles Long-Horizon Robotic Tasks(2 posts)→
More from Embodied
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- OpenRoboto launches decentralized rival to Figure's robot data network on Bittensor — markjeffrey · 2026-09-23
- Typesafe's decision model Jev drives a Go2 robot inside NVIDIA Isaac Sim via OM1 — paigeinsf · 2026-09-23
- Tesla pitches Optimus as the first general-purpose humanoid robot — XFreeze · 2026-09-23
- XPENG's IRON robot demos full-duplex speech with 9-mic array and lip reading — ChrisGPT · 2026-09-23