OpenAI Researcher Lists LLM Shortcomings: Real-Time Voice, Context, Full Multimodality
OpenAI researcher Will DePue laid out current LLM limitations to Nat Friedman: no true real-time barge-in for voice, context pollution in long conversations, missing video encoders and expensive perception, plus a lack of genuinely multimodal output—GPT voice can't do sound effects or beatbox—and odd knowledge gaps.
2026-09-22 ~ 2026-09-22 · 3 related posts
- OpenAI researcher lists LLM bottlenecks: context poisoning, missing video encoders, costly perception — willdepue · 2026-09-22
- OpenAI researcher: LLMs lack true omni output and have weird knowledge holes — willdepue · 2026-09-22
- OpenAI researcher Will Depue on why voice models still lack true realtime chat — willdepue · 2026-09-22