Realtime multimodal models plus steering could yield qualitatively different agents
jakedahn · x · 2026-09-25
Developer jakedahn speculates that once open "jev flavored" models support images, realtime audio and video, you can use them to "surf reality" and respond to the physical world in ways today's text-based agents can't. Pairing that realtime steering with smarter LLMs, he argues, could produce agents that feel qualitatively different — and possibly a path to more fluid agent collaboration.
More from AGI Musings
- 'Put the Fries in the Bag, Terence Tao': A Provocative Essay on Human Dignity in the AI Age — granawkins · 2026-09-25
- Agents can't infer your preferences from a few questions — users must teach the model — jamesbrandecon · 2026-09-25
- Framing AI doomerism as mere neuroticism gives up on predicting the future entirely — AndyMasley · 2026-09-25
- Andy Masley: Framing doomerism as mere neuroticism incapacitates AI debate — AndyMasley · 2026-09-25
- Putting God in the Machine: AI as the next tool for coordinating strangers — imrsn · 2026-09-25
- No AI-proof careers, only adaptable people — plus Paul Graham's watchmaker bet — bendee983 · 2026-09-25