OpenAI Launches Three Real-Time Voice Models
Published · May 8 · Fri Source · openAI (CN)

OpenAI Launches Three Real-Time Voice Models

OpenAI has launched three real-time voice models: GPT-Realtime-2 features GPT-5 level reasoning and tool calling capabilities; GPT-Realtime-Translate supports real-time mutual translation for over 70 languages, costing only about 0.25 yuan per minute, reducing costs by a hundredfold compared to human simultaneous interpretation; GPT-Realtime-Whisper achieves low-latency speech transcription. All three models are available via the Realtime API, with end-to-end processing that preserves tone and emotion.

KeywordsOpenAIGPTAPILaunchesThreeReal-TimeVoiceModels

OpenAI has launched three real-time voice models: GPT-Realtime-2 features GPT-5 level reasoning and tool calling capabilities; GPT-Realtime-Translate supports real-time mutual translation for over 70 languages, costing only about 0.25 yuan per minute, reducing costs by a hundredfold compared to human simultaneous interpretation; GPT-Realtime-Whisper achieves low-latency speech transcription. All three models are available via the Realtime API, with end-to-end processing that preserves tone and emotion.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.