If both VLA and WAM are wrong: three first principles for embodied AI training

qinzytech · x · 2026-09-08

The author lays out three first principles to judge the VLA vs world-model (WAM) debate in embodied AI: (1) superintelligence requires joint training of physical and digital agents, not one or the other, for a 1+1>2 effect; (2) text tokens remain the most efficient representation for reasoning; (3) action and cognition should not be decoupled and should live in the same model. By this rubric, WAM violates each principle by half, VLA's core flaw is violating principle one—and the approach that satisfies all three is "right in front of us," hinting at unified physical+digital joint training grounded in text-token reasoning.

Original post →

More from AGI Musings

AGI Musings channel →