Multimodal Voice Capabilities Remain Underrated

emollick · x · 2026-07-14

The author argues that image input might be one of the most critical capabilities of current models, and tool calling can somewhat substitute for the output abilities of "non-omni" models. In contrast, multimodal voice capabilities seem underexplored, with OpenAI currently driving most of the progress.

Replies note that true "any-to-any" multimodal models haven't become a major industry focus. Google is among the few labs consistently releasing such models; OpenAI is taking a selective multimodal route; Anthropic still lacks multimodal output; and open-weight models show mixed performance.

Related event: Why True Any-to-Any Multimodal Models Haven't Gone Mainstream(3 posts)→

Original post →

More from Models

Models channel →