Alibaba Tongyi Launches Real-time Simultaneous Interpretation Model Qwen3.5-LiveTranslate
Alibaba Tongyi Qianwen launches the real-time simultaneous interpretation model Qwen3.5-LiveTranslate, supporting audio input and text output for 60 languages and audio output for 29 languages, with end-to-end average character latency as low as 2.8 seconds. The model is based on the Qwen3.5-Omni architecture, features real-time voice cloning capabilities to preserve the speaker's original voice characteristics, and includes industry terminology translation optimization with 1,000 built-in hot words.
Alibaba Tongyi Qianwen launches the real-time simultaneous interpretation model Qwen3.5-LiveTranslate, supporting audio input and text output for 60 languages and audio output for 29 languages, with end-to-end average character latency as low as 2.8 seconds. The model is based on the Qwen3.5-Omni architecture, features real-time voice cloning capabilities to preserve the speaker's original voice characteristics, and includes industry terminology translation optimization with 1,000 built-in hot words.
2 days ago • Real-time version of Qwen3.5-LiveTranslate-Flash, a high-precision, high-responsiveness, and high-robustness multilingual real-time audio-video simultaneous interpretation large model. Relying on Qwen3.5-Omni's powerful base capabilities, massive multimodal data, cross-language cross-modal alignment, and visual enhancement...
May 20, 2026 • What is Qwen3.5-LiveTranslate Qwen3.5-LiveTranslate-Flash is a new generation multilingual real-time audio-video simultaneous interpretation model released by the Alibaba Cloud Tongyi Qianwen team, built on the Qwen3.5-Omni Thinker-Talker architecture.
June 12, 2026 • qwen3.5-livetranslate-flash-realtime is a vision-enhanced real-time translation model, supporting mutual translation of 60 languages (29 support audio + text output, 31 support text output only), capable of processing audio and image input simultaneously, suitable for real...
Qwen3.5-LiveTranslate is a real-time simultaneous interpretation large model launched by the Alibaba Tongyi team, supporting input for 60 languages, output for 29 languages, and 3,500+ translation combinations, compressing end-to-end average character latency to 2.8 seconds through readable unit streaming technology, the model features real-time voice cloning and hot...
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.