Multimodal Voice Capabilities Remain Underrated
emollick · x · 2026-07-14
The author argues that image input might be one of the most critical capabilities of current models, and tool calling can somewhat substitute for the output abilities of "non-omni" models. In contrast, multimodal voice capabilities seem underexplored, with OpenAI currently driving most of the progress.
Replies note that true "any-to-any" multimodal models haven't become a major industry focus. Google is among the few labs consistently releasing such models; OpenAI is taking a selective multimodal route; Anthropic still lacks multimodal output; and open-weight models show mixed performance.
Related event: Why True Any-to-Any Multimodal Models Haven't Gone Mainstream(3 posts)→
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22