
OpenAI Expands Real-Time Voice Model Capabilities: GPT-Live-1 Launches on API, Supporting Interruption Handling, Tool Calls, and Telephony Voice Agents
IT Home reported on September 11 that OpenAI today announced the introduction of GPT-Live-1 to its API, enabling developers to build applications and business processes that support real-time voice interaction. The model supports full-duplex conversations, simultaneously listening to and outputting speech, and handling interruptions, pauses, and background noise during conversations. GPT-Live-1 also supports telephony scenarios and can be used for full-duplex voice agents such as restaurant reservations and customer service. OpenAI stated that developers can connect it with tools like Codex, and can also use OpenAI Presence to develop voice workflows that can query enterprise systems, perform approved operations, and transfer to human agents when necessary. According to reports, GPT-Live-1 integrates speech understanding and speech output into the same model, reducing the traditional
Key Takeaways
- Key Highlight:IT Home reported on September 11 that OpenAI today announced the introduction of GPT-Live-1 to its API, enabling developers to build applications and business processes that support real-time voice interaction. The model supports full-duplex conversations, simultaneously listening to and outputting speech, and handling interruptions, pauses, and background noise during conversations. GPT-Live-1 also supports telephony scenarios and can be used for full-duplex voice agents such as restaurant reservations and customer service. OpenAI stated that developers can connect it with tools like Codex, and can also use OpenAI Presence to develop voice workflows that can query enterprise systems, perform approved operations, and transfer to human agents when necessary. According to reports, GPT-Live-1 integrates speech understanding and speech output into the same model, reducing the traditional
- Innovation & Tech:Highlights advancements in OpenAI, GPT, API, demonstrating rapid progress in model capabilities.
- Industry Impact:Reported via IT之家 (CN), offering actionable signals for developers and technology leaders.
IT Home reported on September 11 that OpenAI today announced the introduction of GPT-Live-1 to its API, enabling developers to build applications and business processes that support real-time voice interaction. The model supports full-duplex conversations, simultaneously listening to and outputting speech, and handling interruptions, pauses, and background noise during conversations.
GPT-Live-1 also supports telephony scenarios, and can be used for full-duplex voice agents such as restaurant reservations and customer service. OpenAI stated that developers can connect it with tools like Codex, and can also use OpenAI Presence to develop voice workflows that can query enterprise systems, perform approved operations, and hand over to human agents when necessary.
According to reports, GPT-Live-1 integrates speech understanding and speech output into a single model, reducing the multiple handoffs of the traditional "speech-to-text — large language model — text-to-speech" pipeline, thereby lowering latency. The model also supports offloading complex reasoning and tool calls to a backend text model.
Developers can adjust its tone, speaking rate, and conversation style through system prompts, and can also configure the backend model, tools, and agent framework.
OpenAI stated that GPT-Live-1 supports native speech recognition transcription and reply text, and possesses alphanumeric understanding, keyword biasing, and turn detection capabilities.
In early evaluations, after the language learning platform Speak used GPT-Live-1, the number of erroneous interruptions during learners' thinking pauses decreased by nearly 80% compared with the previous turn-based system. Tony Stoyanov, co-founder and CTO of a healthcare services platform, said that his team reduced its codebase by 80%, deleting approximately 23,000 lines of code.
Evaluation results released by OpenAI show that GPT-Live-1 outperformed GPT-Realtime-2.1 by 30 percentage points on the Full Duplex Bench test, and ranked first on the Tau3 end-to-end voice agent task benchmark when paired with GPT-6 Astra at a medium reasoning intensity.
Currently, GPT-Live-1 is available in the API, with the front-end voice layer priced at $0.05 per minute. Calculated at 60 minutes per hour, the cost of the front-end voice layer alone is approximately $3 (IT Home note: equivalent to about 20.2 RMB at the current exchange rate). Backend model and agent tool costs are billed separately.
OpenAI also expanded its real-time voice options, adding voices such as Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, and Cinder, and plans to continue adding voice and language support in the coming months.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.
Industry Insights & Analysis
As artificial intelligence rapidly evolves, breakthroughs surrounding OpenAI, GPT, API, Agent are shifting toward scalable, robust real-world implementations.
Driven by both open-source ecosystems and proprietary model architectures, the integration between compute optimization, data engineering, and agentic workflows is accelerating. This development provides a strategic benchmark for upcoming AI tooling and developer workflows.