Explicit gesture input failed AR/VR — AI contextual interpretation is the missing piece

mrjonfinger · x · 2026-09-06

Jon Finger argues investors distrust AR/VR because explicit gesture input failed in the last hype cycle — it demands exhausting, precise repetition from users.

The real missing element, he says, is fully interpretive gesture understanding grounded in visual context — the same non-language-specific interpretation power diffusion models showed with text-to-image. Once that works, AR/VR changes entirely, and natural language recedes to casual social input, at best the third-most-important input to physical action.

The quoted thread adds: AI now makes building 3D worlds cheap, but the rest — hardware, a proper OS, UI/UX, real 6DoF experiences — is still missing.

Original post →

More from AGI Musings

AGI Musings channel →