If both VLA and WAM are wrong: three first principles for embodied AI training
qinzytech · x · 2026-09-08
The author lays out three first principles to judge the VLA vs world-model (WAM) debate in embodied AI: (1) superintelligence requires joint training of physical and digital agents, not one or the other, for a 1+1>2 effect; (2) text tokens remain the most efficient representation for reasoning; (3) action and cognition should not be decoupled and should live in the same model. By this rubric, WAM violates each principle by half, VLA's core flaw is violating principle one—and the approach that satisfies all three is "right in front of us," hinting at unified physical+digital joint training grounded in text-token reasoning.
More from AGI Musings
- 89% of firms deploy AI but only 6% see real returns, and AI layoffs hit 112,713 — mikeflache · 2026-09-08
- Critic says AI doubters quietly pivoted from 'snake oil' to 'normal technology' after being proven wrong — nitarshan · 2026-09-08
- Astra Driving Blender via Computer Use Could Be AI's Fourth Demand Wave — firstadopter · 2026-09-08
- Jack Clark publishes 'The Thousand and One Faces of Repair,' fiction about fixing misbehaving AIs under 2030s sentience accords — jackclarkSF · 2026-09-08
- Is AI slop okay at work? Debate over raw AI output in communication — oliviaakory · 2026-09-08
- Steven Pinker: why AI won't kill us—and his take on the value alignment problem — davidmanheim · 2026-09-08