Realtime multimodal models plus steering could yield qualitatively different agents

jakedahn · x · 2026-09-25

Developer jakedahn speculates that once open "jev flavored" models support images, realtime audio and video, you can use them to "surf reality" and respond to the physical world in ways today's text-based agents can't. Pairing that realtime steering with smarter LLMs, he argues, could produce agents that feel qualitatively different — and possibly a path to more fluid agent collaboration.

Original post →

More from AGI Musings

AGI Musings channel →