N_0-VTLA: First VTLA Foundation Model Pretrained on Tactile Data at Scale
NeoteAIEmbodied · hf · 2026-08-03
Researchers introduced N0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of fine-grained contact-rich manipulation. It is the first VTLA model pretrained on tactile data at scale.
Core Technical Highlights:
- Training Recipe: Features visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement.
- Data: Learns broad contact priors from NeoData, a large-scale visuo-tactile robot dataset.
- ALTER (Offline RL): Converts relative progress and trajectory events into binary advantage labels, significantly boosting contact-rich skills like deformable object manipulation.
Performance:
- Wins all 9 real-robot NeoReal tasks.
- Achieves 63.8% mean success on a 20-task simulation suite (vs. 44.0% for the strongest baseline).
- Reaches 75-95% success on three long-horizon real-robot tasks.
More from Embodied
- AgenticROS Launches Cloud Service with Global P2P Teleop and Multi-Hardware Support — chrismatthieu · 2026-08-03
- RoboHarness: Heterogeneous Policy Orchestration Boosts Long-Horizon Robotics to 95.2% — 机器之心 · 2026-08-03
- Foundation Robot Catches Baseballs with Open-Loop Control and Low-Friction Tendons — CyberRobooo · 2026-08-03
- Jensen Huang: Video Generation Models Have Solved the Physics Intuition for Robotics — r0ck3t23 · 2026-08-03
- OpenAI Hardware Roadmap Leaks: $200 Screenless Speaker and AI Phone Target 2027 — 创业邦 · 2026-08-03
- ACE-Data-0: New Embodied Data Engine on Hugging Face with 150 Hours of Multimodal Data — liuziwei7 · 2026-08-03