VLA Reading Room: World Models Need Co-Training with Actions
agihouse_org · x · 2026-08-28
AGI House and Radical Ventures hosted a deep-dive discussion on Vision-Language-Action (VLA) models, led by Vincent Vanhoucke, co-author of RT-2. Key takeaways include:
- World models and action prediction must be co-trained: Pure action training misses physical规律; pure video leads to loose, hacky physics. Co-training forces the model to focus on what matters, improving performance. DreamZero hints at this but lacks the ablation study.
- Hard-coding physics constraints degrades performance at scale: The "Bitter Lesson" applies to robotics; encoding priors too precisely (e.g., friction coefficients) limits model potential as it scales.
More from Embodied
- Isaac 0.5 robot plays chess: perception, decision, actuation — code_star · 2026-08-28
- Forecast: 1 Billion Optimus Robots by 2036, Best-Selling Product Ever — tetsuoai · 2026-08-28
- BeyondMimic enables humanoid robots to perform aerial cartwheels — kevin_zakka · 2026-08-28
- Microduck BAM Structure Analysis: Exploring the Feasibility of Adding Actuators — _Stocko_ · 2026-08-28
- Cortisol comparison: Waymo induces lower stress than Uber rides — beffjezos · 2026-08-28
- Claude Controls Lab Gear: Laser Fix Rate Hits 99.3% in Tests — chrismattmann · 2026-08-28