OpenAI Researcher Lists LLM Shortcomings: Real-Time Voice, Context, Full Multimodality

OpenAI researcher Will DePue laid out current LLM limitations to Nat Friedman: no true real-time barge-in for voice, context pollution in long conversations, missing video encoders and expensive perception, plus a lack of genuinely multimodal output—GPT voice can't do sound effects or beatbox—and odd knowledge gaps.

2026-09-22 ~ 2026-09-22 · 3 related posts