Google Launches Real-Time Translation Model Gemini 3.5 Live Translate
Published · Jun 10 · Wed Source · IT之家 (CN)

Google Launches Real-Time Translation Model Gemini 3.5 Live Translate

Google officially launches the real-time voice translation model Gemini 3.5 Live Translate, which can automatically recognize over 70 languages and generate fluent translation audio that preserves the original speaker's tone, speed, and pitch. The model balances waiting for context with instant translation, lagging only a few seconds behind the speaker. The model is open to developers via the Gemini Live API preview, enterprise users can use it in Google Meet, and general users can experience it through the Google Translate App.

KeywordsGoogleGeminiAPILaunchesReal-TimeTranslationModelLive

Twenty years ago, Google Translate officially started as one of the pioneer experiments in the field of machine learning, dedicated to transforming language science into a bridge for communication between people. After years of development, this project has covered billions of users, translating over one trillion words per month.

Today, Google reaches a new milestone, officially releasing Gemini 3.5 Live Translate—Google's latest audio model, designed specifically for real-time speech-to-speech translation.

Currently, Gemini 3.5 Live Translate is beginning to roll out across multiple Google products:

For developers: Open public preview via Gemini Live API and Google AI Studio;

For enterprise users: Launched as a private preview in Google Meet starting this month;

For everyone: Officially launched in the Android and iOS versions of the Google Translate app.

Developer Integration with Gemini 3.5 Live Translate

Through the Gemini Live API, developers can achieve video dubbing and multi-language synchronous translation. Interested developers can visit the Gemini Cookbook to view demo examples and more reference code.

Developer platforms such as Agora, Fishjam, LiveKit, Pipecat, and Vision Agents have integrated the Gemini Live API, helping developers build and deploy voice translation applications more conveniently. These integration solutions encapsulate complex real-time media streaming infrastructure, allowing developers to focus on the user experience itself.

Partner Grab is testing the model to achieve near real-time multi-language communication between drivers and passengers. Grab platform users generate over 10 million communication needs via voice calls per month.

Positive Feedback from Partners

Besides Grab, companies such as CJ ENM and LiveKit have also given positive evaluations of Gemini 3.5 Live Translate, generally believing its translation quality is excellent, accuracy is high, and latency performance is outstanding.

Experience Real-Time Translation Upgrade in Google Meet

Google Meet's voice translation feature will soon integrate Gemini 3.5 Live Translate, bringing the following improvements:

Supported languages expanded from only 5 previously to over 70;

A single meeting can support translation between over 2000 language combinations, no longer limited to English as the intermediate language;

Interface updated, users can instantly call the voice translation feature.

This update will be open to some Google Workspace enterprise users as a private preview starting this month, and will be launched to more users later this year.

Experience Gemini 3.5 Live Translate in the Google Translate App

The model is simultaneously launched globally in the Android and iOS versions of the Google Translate app. When using the real-time translation feature, simply connect any headphones to experience fluent translation supporting over 70 languages while preserving the speaker's tone characteristics.

For Android users, Google is also launching a new "Earpiece Mode," allowing users to receive translation audio through the earpiece without wearing headphones, just by holding the phone close to the ear like answering a normal call. This mode is suitable for scenarios where it is inconvenient for others to hear the translation content, providing a more private and convenient user experience. For example, users can receive real-time English translation via earpiece mode while listening to a Spanish tour guide.

SynthID Watermark Ensures Security

Q&A

Q1: What is the difference between Gemini 3.5 Live Translate and traditional translation systems?

A: Gemini 3.5 Live Translate uses a continuous generation method, without needing to wait for the speaker to finish a full sentence before translating. It can follow the speaking rhythm in real-time while ensuring quality, with almost no obvious pauses throughout, maintaining synchronization within seconds of the speaker. Traditional sentence-by-sentence translation systems need to wait until the complete statement ends before starting translation, resulting in stronger latency and lower naturalness.

Q2: Which languages does Gemini 3.5 Live Translate support, and how do general users use it?

Q3: Is the audio content generated by Gemini 3.5 Live Translate safe, and are there mechanisms to prevent abuse? Return to Sohu, view more.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.