Explicit gesture input failed AR/VR — AI contextual interpretation is the missing piece
mrjonfinger · x · 2026-09-06
Jon Finger argues investors distrust AR/VR because explicit gesture input failed in the last hype cycle — it demands exhausting, precise repetition from users.
The real missing element, he says, is fully interpretive gesture understanding grounded in visual context — the same non-language-specific interpretation power diffusion models showed with text-to-image. Once that works, AR/VR changes entirely, and natural language recedes to casual social input, at best the third-most-important input to physical action.
The quoted thread adds: AI now makes building 3D worlds cheap, but the rest — hardware, a proper OS, UI/UX, real 6DoF experiences — is still missing.
More from AGI Musings
- AI has created around 1M new jobs in America, Economist analysis finds — pmddomingos · 2026-09-06
- Brundage: policymakers will eventually freak out that AI labs can't defend their IP — Miles_Brundage · 2026-09-06
- Dev warns hand-coding features without following AI is 'sitting on a beach as a tsunami approaches' — draginol · 2026-09-06
- Harvard-led arXiv paper models LLM adoption as a 'cognitive virus' with dependence tipping points — rohanpaul_ai · 2026-09-06
- Harvard-Led Paper Warns Small LLM Adoption Rises Could Trigger Cognitive Dependence Tipping Point — rohanpaul_ai · 2026-09-06
- Danielle Fong: in the singularity, get rich by being first, smarter, or cheating — ZeroStateReflex · 2026-09-06