OpenAI Launches New Generation Voice Model GPT-Live
Published · Jul 9 · Thu Source · 量子位 (CN)

OpenAI Launches New Generation Voice Model GPT-Live

OpenAI has launched a new voice interaction model, GPT-Live, supporting real-time simultaneous interpretation and full-duplex conversation, allowing users to interrupt at any time. The model supports customizable reasoning intensity and introduces a deep task delegation mechanism. While maintaining smooth voice interaction in the foreground, it can process tasks such as web search and complex reasoning in parallel in the background. Visual cards for weather, events, etc., can pop up in real-time during conversations.

KeywordsOpenAIGPTLaunchesNewGenerationVoiceModelGPT-Live

OpenAI Releases GPT-Live Late at Night, ChatGPT Finally Speaks Like a Real Person

Late at night, OpenAI suddenly launched GPT-Live.

This new generation voice model pushes ChatGPT Voice a step forward: human-AI communication begins to resemble a real conversation.

GPT-Live is built on a full-duplex architecture, meaning it can listen and speak simultaneously. During conversation, GPT-Live can indicate it is listening with responses like "mhmm" or "yeah," and engage in rapid back-and-forth exchange; when you need time to think, it remains quiet. Ultimately, this delivers a very relaxed and natural voice interaction experience.

OpenAI states that GPT-Live is their most intelligent voice model to date. For questions requiring web search, deeper reasoning, or more complex processing, it delegates tasks to this latest frontier model in the background and brings the results back into the conversation once ready. While handling these tasks, GPT-Live can continue communicating with you, maintaining conversation flow.

At launch, the model called by GPT-Live in the background is GPT-5.5. As OpenAI releases new frontier models subsequently, the background model called by GPT-Live will continue to update.

Starting today, OpenAI will gradually roll out two versions of GPT-Live to ChatGPT users globally: GPT-Live-1 and GPT-Live-1 mini.

Previous Voice AI Roadmap

Before GPT-Live, the previous generation voice AI system had already advanced human-machine voice interaction towards natural conversation, but different technical roadmaps had obvious trade-offs.

One is the cascaded voice system.

Early voice systems typically used a cascaded architecture, where multiple models completed one round of conversation processing sequentially. The original ChatGPT Voice adopted this method: first, a speech-to-text model transcribed user voice into text, then a large language model generated a reply, and finally, a text-to-speech model converted the reply into voice.

This architecture allowed users to communicate with frontier AI models via voice for the first time, but the series connection of multiple models brought extra complexity. Information could be lost when passing between different models, and system responses tended to appear slow and stiff.

Turn-based voice models.

Subsequent turn-based voice models, such as ChatGPT Advanced Voice Mode, began processing and generating audio within a single model, reducing latency and making the conversation experience smoother.

However, such models still followed a one-turn-back-and-forth mechanism. The model needed to wait for the user to stop speaking before responding, so the interaction method remained relatively fixed. Meanwhile, since the system typically relied on silence to judge whether a turn of speech had ended, pauses during brief user thinking or background noise in the environment could be misjudged as the end of speech, causing the model to interrupt at unnatural times.

GPT-Live New Method

GPT-Live addresses the above limitations through two architectural changes.

Continuous Interaction.

First, GPT-Live is based on a full-duplex architecture, designed for continuous interaction. It does not process individual independent messages, but continuously processes input while generating output. Therefore, the model can make interaction decisions multiple times per second, including whether to speak, continue listening, pause, interrupt, or call tools.

This allows the model to engage in more natural back-and-forth exchange, better grasp the timing rhythm, and even complete real-time translation.

Delegation Mechanism for Deep Tasks.

Second, OpenAI decouples GPT-Live, responsible for continuous interaction, from deeper task processing capabilities. When a question requires search, reasoning, or stronger agent capabilities, GPT-Live can delegate the task to other models like GPT-5.5. This allows it to continue maintaining the conversation even while processing multiple tasks in the background.

This architectural change also allows GPT-Live to continuously use the latest models and agents, combining frontier intelligence with natural interaction.

Evaluation

OpenAI built a new human evaluation system to measure conversation pleasantness and fluency. In these one-to-one comparison tests, GPT-Live-1 and GPT-Live-1 mini were significantly preferred over Advanced Voice Mode in 5 to 10 minute conversations under matching conditions. Evaluation dimensions included overall preference, turn transition, interruption situations, conversation fluency, and the naturalness of each interaction.

GPQA: GPT-Live-1 significantly outperforms Advanced Voice Mode on GPQA. GPQA is used to test models' expert-level scientific reasoning capabilities in fields such as biology, chemistry, and physics.

BrowseComp: GPT-Live-1 achieves significant improvement over Advanced Voice Mode on BrowseComp. BrowseComp is used to test agent-style web search capabilities and the model's ability to locate hard-to-find information.

τ³-Voice Telecom (Internal Variant): GPT-Live-1 outperforms Advanced Voice Mode on τ³-Voice Telecom. This evaluation tests the performance of voice agents in real, multi-turn telecom customer service tasks.

Brand New ChatGPT Voice Experience

More than 150 million people converse with ChatGPT weekly through features like Voice and Dictation. They use it for hands-free daily assistance, language practice, bedtime stories, or just casual chat during commutes.

Starting today, when you click the Voice button to converse with ChatGPT, you will receive an upgraded experience driven by GPT-Live, including more natural conversation, smarter answers, better listening capabilities, and visual replies.

More Natural Conversation

Now, conversing with ChatGPT will be closer to real communication. You can interrupt to ask questions at any time, stop to organize your thoughts, or ask ChatGPT to slow down its speech rate. It will naturally indicate it is following your words with responses like "mhmm" or "got it," letting you know it is always listening. The team has also remastered nine different voices in ChatGPT for GPT-Live.

Smarter Answers

ChatGPT Voice can now call OpenAI's latest frontier models to provide smarter answers when needed. You can also choose different reasoning intensities based on needs: Instant is suitable for quick responses, while Medium and High are suitable for scenarios where you want ChatGPT to spend more time thinking.

Better Listening Capabilities

When you need a little time to think, ChatGPT Voice will now wait instead of rushing to interrupt. If you ask it to stay quiet and continue listening, it will do so. When there is background noise around, such as vehicles passing or people talking nearby, ChatGPT can focus better on your voice and is less likely to be distracted.

Visual Answers Visible at a Glance

Some answers are more useful when visible. During voice conversations, ChatGPT can now display rich visual cards for topics such as weather, stocks, and sports. Voice will continue to support search, memory, images, and file uploads.

Ultimately, ChatGPT Voice becomes more natural, more capable, and better suited for various usage scenarios in daily life.

Finally, GPT-Live is currently rolling out to global ChatGPT users, covering iOS, Android, and ChatGPT.com. GPT-Live-1 will become the default model for Go, Plus, and Pro users when using ChatGPT Voice, and GPT-Live-1 mini will become the default model for Free users.

Reference Link: https://openai.com/index/introducing-gpt-live/

This article comes from the WeChat public account "Synced", published by 36Kr with authorization.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.