FULL STORY

OpenAI Voice: From Launch to Cross-Platform Tests

OpenAI launched GPT-Live and expanded it to desktop and iOS. Users praise its natural interaction and cross-platform capabilities, calling it a major step toward AGI.

2026-07-07 ~ 2026-08-01 · 12 episodes · 177 posts

Episode 1 · OpenAI Launches Full-Duplex Voice Model GPT-Live (2026-07-07, 44 posts)

On July 9, OpenAI officially launched GPT-Live, a new full-duplex voice model now available in ChatGPT, marking a major upgrade in AI voice interaction. The model supports natural, simultaneous listening and speaking with the ability to be interrupted anytime.

Key Details and Versioning

A core breakthrough of GPT-Live is the separation of voice conversation from heavy computational tasks. When encountering complex problems requiring search or deep reasoning, the model offloads the task to a frontier model (GPT-5.5 at launch) in the background while keeping the voice conversation uninterrupted. Additionally, it can generate real-time UI interfaces, displaying information via visual cards for weather, stocks, and sports. Regarding versions, GPT-Live-1 serves as the default for Go, Plus, and Pro users, while GPT-Live-1 mini is available for free users. OpenAI also previewed API access for both versions and opened notification registrations for developers and enterprises.

Performance and Product Evolution

According to data shared by @testingcatalog, GPT-Live-1 significantly outperforms the previous Advanced Voice Mode across multiple dimensions, including GPQA, BrowseComp, and internal τ³-Voice Telecom tests. The model was initially known as "GPT Bidi 1" before being officially renamed GPT Live 1 to unify the voice product line's naming convention. The company also showcased new capabilities like image interaction during recent live demonstrations.

24 more related posts →

Episode 2 · New ChatGPT Voice Mode Tested: Near-Human Multi-lingual Experience (2026-07-09, 14 posts)

Recently, ChatGPT's voice mode received a major update. After testing the latest app version, multiple users reported a significant improvement in the voice interaction experience, with greatly enhanced realism and fluidity, even evoking a sci-fi sense of being close to AGI. This marks a shift in AI voice interaction, breaking away from traditional turn-based Q&A paradigms to become more natural.

Multi-lingual and Realistic Performance

Many users noted that the new voice mode performs exceptionally well in non-English contexts. @asianteaman was amazed by the absurd realism in Russian, where the model naturally pauses, sighs, randomly giggles, and even acts tired, contrasting with the more customer-service-like English voice. @SteeeeveJune and @Bob-the-Human praised the natural German pronunciation and the realistic, thinking-aloud pauses. @flowersslop added that accents vary by language, making the overall experience smooth, and specifically liked the cute voice of Maple.

Interaction and Emotional Feedback

The new version demonstrates greater intelligence and emotional resonance. @Healthy-Nebula-3603 mentioned the model performs better during web searches, even proactively asking for quiet and closing the conversation itself when finished. @flowersslop had a 40-minute continuous chat about personal emotional struggles, experiencing genuine empathy and active listening. Both @Dimillian and @RileyRalmuto suggested that this natural interaction paradigm is ending traditional Q&A formats.

Controversies and Limitations

Despite the overwhelmingly positive feedback, boundaries remain. @emollick specifically reminded users that the voice model does not equate to a full reasoning model, and its limitations must be kept in mind. Additionally, @Bob-the-Human pointed out that the new voices have a lower volume, making them hard to hear in noisy environments like cars, and noted that previously available French or German accents might now be restricted.

Episode 3 · ChatGPT's New Voice Mode Receives Positive Feedback (2026-07-11, 3 posts)

ChatGPT's new voice mode is praised as a major step forward, offering a more natural interaction that could reduce screen time. However, it still has minor limitations, such as struggling with simple letter counting.

Episode 4 · GPT-Live Voice Mode Receives Praise for Major Upgrades (2026-07-13, 3 posts)

Users are reporting highly positive experiences with the upgraded GPT-Live voice mode. The new system offers millisecond-level latency, natural conversations, and excellent emotional expression in real-time translation, despite minor limitations like lacking distinct voices for different speakers.

Episode 5 · OpenAI Launches GPT-Live for Real-Time Full-Duplex Voice Interaction (2026-07-13, 6 posts)

In mid-July, OpenAI officially launched GPT-Live, its next-generation voice model. The core focus is to push voice interaction towards a more natural, real-time "full-duplex" conversational mode. This release marks a shift for AI voice capabilities from traditional turn-based interactions to fluid, real-time communication, drawing significant attention from the community.

Key Details

Official materials position GPT-Live as a major upgrade for voice capabilities. OpenAI released an introductory page and a demo video titled "Background Robustness with GPT-Live," highlighting the model's ability to maintain solid voice interaction performance even in environments with background noise or ambient sound.

External Interpretations and Tests

Several creators shared more specific functional insights following the launch. emmanuelvivier pointed out that GPT-Live features real-time full-duplex voice, allowing users to interrupt and speak naturally without waiting for the model to finish its turn. thione added that the model supports more natural interruptions, real-time translation, and real-time conversation. Furthermore, thursdaipod noted in testing that OpenAI's newly shipped full-duplex conversational engine enables genuinely natural back-and-forth speaking, marking a significant milestone since the Scarlett Johansson voice controversy.

Episode 6 · GPT Live Voice Praised for Speed, Criticized for Accent Gaps (2026-07-16, 5 posts)

Multiple users shared hands-on impressions of OpenAI’s GPT Live real-time voice mode this week, and the feedback was notably mixed. Positive reports focused on stability, speed, and the ability to complete practical tasks inside a voice session, while criticism centered on accent handling and weaker multilingual performance. Taken together, the posts suggest GPT Live is improving in core usability, but still delivers uneven value across languages.

Experience gains

@IntelligentKey2760 said that after using GPT-Live for a week, it felt more stable than the older voice mode. In the author’s comparison, the previous mode could sometimes respond instantly and sometimes pause for several seconds, making the rhythm feel like an unreliable long-distance call; GPT-Live, by contrast, maintained a more consistent cadence with almost no long stalls. The author also noted that on complex questions, it sometimes uses fillers such as “um” or “emmm,” which makes the exchange feel more human.

@D3VAUX also gave a favorable test result, saying the real-time voice mode felt significantly faster and could combine live data retrieval with UI component display. As an example, the system reportedly used weather conditions and user preferences to plan an evening activity within 5 minutes, while also checking pricing from different merchants.

Main shortcomings

@Angaisb argued that OpenAI should not have shipped GPT Live before polishing accent support and multilingual performance. In the author’s view, the feature feels “like magic” for English users, but non-English users are not getting the same level of experience, and that gap has already lasted for months.

In a follow-up, @Angaisb added that GPT Live is indeed good in some respects, but still requires manual prompting and is not yet as natural as the original AVM. The author expressed hope that GPT Live 1.1 or 1.5 will address these issues.

Episode 7 · ChatGPT's New Voice Mode Wins User Praise (2026-07-22, 2 posts)

ChatGPT's new GPT-Live voice mode is receiving high praise for its lifelike interactions. Users report that the feature delivers an impressive, natural conversational experience with realistic breathing pauses and subtle responses.

Episode 8 · ChatGPT Voice Hits Desktop and iOS as Voice Control Hub (2026-07-23, 80 posts)

OpenAI has officially brought ChatGPT Voice to its desktop apps (macOS and Windows) and the iOS Remote app, available to Plus, Pro, Business, Edu, and Enterprise plan users. Powered by GPT-Live, the feature goes beyond voice chatting, allowing users to control their computers and coordinate multiple agents within ChatGPT Work and Codex. This marks ChatGPT's transition into a voice hub capable of operating the system and orchestrating background agents.

Confirmed

ChatGPT Voice allows users to delegate tasks via voice and have the computer execute them automatically. It supports coordinating tasks across Chat, Work, and Codex, scheduling multiple running agents. The desktop update also includes support for taking app screenshots, using the browser, interacting with MCPs, continuing work across project threads, and mounting multiple folders for local projects. Additionally, the iOS Remote app now supports voice control, enabling users to use their phones as hands-free controllers to talk to their computers and operate Codex (described as a "Codex phone"). Official confirmation notes that Android support is coming soon, and the feature is currently rolling out to more users.

Real-World Experience and Feedback

Several developers shared their workflows and evaluations. @Dimillian stated that initiating and managing Codex tasks via voice on mobile or desktop is incredibly fluid, calling it a true AGI moment. @jxnlco pointed out that the voice feature acts as a "planning mode," clarifying ambiguous requirements through follow-up questions before launching Codex tasks. He emphasized that the core mental model for Voice is an "orchestrator" rather than an executor, routing tasks to appropriate threads. @petergyang agreed that it acts as an orchestrator thread in Codex, though users currently sometimes need to manually tell it to "start new thread." @omooretweets added that the voice feature integrates with the visual interface for more natural interaction, directly calling tools to check calendars and schedule events. @reachvb suggested treating it as a "chief of staff" to clarify issues, assign tasks, and track progress.

Why it matters

By integrating desktop and iOS remote control, this update further reduces users' reliance on keyboards and mice, making barrier-free, speak-aloud productivity and parallel development a reality.

60 more related posts →

Episode 9 · ChatGPT Voice Hailed as an 'AGI Moment' (2026-07-24, 3 posts)

ChatGPT Voice is being widely praised by users as an 'AGI moment.' Its ability to understand casual, scattered spoken thoughts and turn them into actionable items is seen as a major upgrade to AI workflows.

Episode 10 · ChatGPT Voice Lands on Desktop with Agentic Control (2026-07-26, 5 posts)

OpenAI has officially brought ChatGPT Voice to macOS and Windows desktops, rolling it out globally to Plus, Pro, Business, Edu, and Enterprise users. This update allows users to direct multi-step tasks and agentic workflows on their computers via voice, marking a major upgrade to ChatGPT's voice interaction capabilities.

Confirmed

Regarding coverage, ChatGPT Voice is confirmed to be available globally to subscribers across Plus, Pro, Business, Edu, and Enterprise tiers. The underlying technology is powered by the new ChatGPT-Live (or GPT-Live) voice model. In terms of interactive experience, the feature focuses on "asymmetric voice interaction," meaning users can input complex ideas via voice, and the model can autonomously decide the best format (not necessarily voice readout) to present the answer. For application scenarios, desktop voice can not only chat but also directly control the computer, integrating with agentic processes like ChatGPT Work and Composer to take over multi-step tasks.

Delayed Rollout Reason

According to @athyuttamre, the rollout for Edu, Business, and Enterprise sectors was slightly delayed compared to individual users because OpenAI needed to refine admin controls and other enterprise-level compliance requirements first.

Why it matters

Deeply integrating voice capabilities with desktop agents signifies a step forward for AI from a mere conversational tool to a "computer operation assistant." Voice input is better suited for expressing complex instructions than typing, and this asymmetric interaction model significantly lowers the barrier for users to manipulate multi-agent workflows.

Episode 11 · ChatGPT Voice Mode Tested: From Voice Assistant to Hands-Free Computer Control (2026-07-29, 10 posts)

Recent in-depth tests by multiple developers and users show that ChatGPT's voice mode has evolved from simple dictation into a new workflow. Especially when combined with Codex, users can not only verbally command AI to perform computer operations but also remotely invoke desktop computing power via phone. This marks a substantial shift in human-computer interaction, with seamless voice dialogue becoming the next core interface.

Confirmed

  • @Dimillian demonstrated a workflow using ChatGPT Voice with Codex, where users verbally instruct AI to select files in Finder and execute AppShot, achieving hands-free computer operation.
  • @dfinke conducted a live demo of ChatGPT Voice in Codex without a fixed script, allowing interruptions, direction changes, research, and continued building, verifying how conversation becomes part of the workflow.
  • @pbbakkum provided a pairing tutorial: add a remote device in desktop paid settings, then use voice commands on mobile to remotely invoke desktop computing power for tasks like creating projects.
  • @rudrank shared experience using voice mode to plan schedules by integrating local files, Slack, and Gmail, comparing its natural interaction to the movie 'Her'.
  • @EverydayAI summarized 7 best ways to use voice mode, likening it to 'Jarvis'. @nickbaumann called it the most impactful AI tool for his workflow.
  • @机器之心 reported a case where a blogger used ChatGPT voice mode for 4 hours while hiking, claiming higher efficiency than 8 hours in the office. Sam Altman retweeted it, exclaiming about a 'new computer'.
  • @danshipper said every member of his team Every was amazed by ChatGPT for Work's voice mode.

Why it matters

  • These tests show AI voice interaction has moved beyond simple Q&A to deeply intervene in system-level software operations. This 'speak and it's done' mode greatly improves efficiency and may reshape basic computer usage habits.

Episode 12 · ChatGPT Desktop Apps Launch New Voice Mode for Cross-App Workflows (2026-07-30, 2 posts)

OpenAI has introduced a new voice interaction mode for ChatGPT on macOS and Windows desktops. This feature allows users to seamlessly control cross-app workflows via voice commands without interrupting their current tasks, enabling a hands-free experience.